2026
HitKV: Activation Frequency Knows Which Tokens Are Important
AAAI 2026technical
The demand for long-context processing in large language models (LLMs) continues to escalate alongside rapid advancements in their capabilities. However, the intermediate attention keys and values (KV cache) employed to avoid re-computations, also grow linearly with sequence length, far exceeding th