2026
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
ICLR 2026poster
Transformer-based large language models (LLMs) rely on key–value (KV) caching to avoid redundant computation during autoregressive inference. While this mechanism greatly improves efficiency, the cache size grows linearly with the input sequence length, quickly becoming a bottleneck for long‑context…