2026
Beyond Speedup - Utilizing KV Cache for Sampling and Reasoning
ICLR 2026poster
KV caches, typically used only to speed up autoregressive decoding, encode contextual information that can be reused for downstream tasks at no extra cost. We propose treating the KV cache as a lightweight representation, eliminating the need to recompute or store full hidden states. Despite being w…