2025
LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
EMNLP 2025
KV Cache is commonly used to accelerate LLM inference with long contexts, yet its high memory demand drives the need for cache compression. Existing compression methods, however, are largely heuristic and lack dynamic budget allocation. To address this limitation, we introduce a unified framework fo