← Search

Aomufei Yuan

1 accepted papers

2026

KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference

AAAI 2026technical

Efficient inference of large language models (LLMs) is hindered by an ever-growing key-value (KV) cache, making KV cache compression a critical research direction. Traditional methods selectively evict less important KV cache entries, which leads to information loss and hallucinations. Recently, mer

Cited by 0SourcePDFScholar