2026
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
AAAI 2026technical
Efficient inference of large language models (LLMs) is hindered by an ever-growing key-value (KV) cache, making KV cache compression a critical research direction. Traditional methods selectively evict less important KV cache entries, which leads to information loss and hallucinations. Recently, mer