← Search

Song Yuan

1 accepted papers

2025

LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation

EMNLP 2025

KV Cache is commonly used to accelerate LLM inference with long contexts, yet its high memory demand drives the need for cache compression. Existing compression methods, however, are largely heuristic and lack dynamic budget allocation. To address this limitation, we introduce a unified framework fo