2026
EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments
ICML 2026poster
Modern large language models (LLMs) extend context lengths to millions of tokens, enabling coherent, personalized responses grounded in long conversational history. However, the Key-Value (KV) cache grows linearly with the extended dialogue history, causing the model’s memory footprint to quickly ex…