2026
Channel-Aware Mixed-Precision Quantization for Efficient Long-Context Inference
ICLR 2026poster
The key-value (KV) cache plays a vital role in accelerating autoregressive inference for large language models (LLMs). However, its linear memory growth with sequence length poses significant memory bottlenecks, especially in long-context scenarios. Quantization offers a promising solution for memor…