2025
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
NeurIPS 2025spotlight
KV cache eviction has emerged as an effective solution to alleviate resource constraints faced by LLMs in long-context scenarios. However, existing token-level eviction methods often overlook two critical aspects: (1) their irreversible eviction strategy fails to adapt to dynamic attention patterns…