CLAA: Cross-Layer Attention Aggregation for Accelerating LLM Prefill
The prefill stage in long-context LLM inference remains a computational bottleneck. Recent token-ranking heuristics accelerate inference by selectively processing a subset of semantically relevant tokens. However, existing methods suffer from unstable token importance estimation, often varying betwe…