2025
CateKV: On Sequential Consistency for Long-Context LLM Inference Acceleration
ICML 2025poster
Large language models (LLMs) have demonstrated strong capabilities in handling long-context tasks, but processing such long contexts remains challenging due to the substantial memory requirements and inference latency. In this work, we discover that certain attention heads exhibit sequential consist…