2026
KVmix: Gradient-Based Layer Importance-Aware Mixed-Precision Quantization for KV Cache
AAAI 2026technical
The high memory demands of the Key-Value (KV) Cache during the inference of Large Language Models (LLMs) severely restrict their deployment in resource-constrained platforms. Quantization can effectively alleviate the memory pressure caused by KV Cache. However, existing methods either rely on stati