2026
AdaCache: Adaptive Caching and Context Augmentation for Efficient LLM Serving
ICLR 2026poster
Retrieval-Augmented Generation (RAG) significantly enhances Large Language Models by integrating external knowledge sources, but at the cost of substantial computational overhead from extended input sequences. Current RAG systems exhibit two fundamental inefficiencies: redundant processing of frequ…