2025
Enhancing Large Language Model Inference Efficiency via Lookahead Cache Filtering
ICASSP 2025accepted
The large Key-Value (KV) cache is a significant challenge in deploying Large Language Models (LLMs). Current research addressing these issues employs cache compression techniques, which we find suffer from information loss and the "lost-in-the-middle" problem. We propose the Lookahead Cache Filterin…