2025
QoKV: Comprehending and Surpassing the Hurdles of KV Cache Quantization
ICASSP 2025accepted
Large language models (LLMs) have demonstrated outstanding performance in various tasks. However, the memory footprint of the key-value (KV) cache generated during model inference poses significant challenges for efficient model deployment. This paper presents a detailed analysis of the KV cache and…