← Search

Wang Xiaoliang

1 accepted papers

2025

SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference

NeurIPS 2025spotlight

KV cache eviction has emerged as an effective solution to alleviate resource constraints faced by LLMs in long-context scenarios. However, existing token-level eviction methods often overlook two critical aspects: (1) their irreversible eviction strategy fails to adapt to dynamic attention patterns…

Cited by 0SourceScholar