2026
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
ICML 2026poster
Large reasoning models achieve strong performance through test-time scaling, but this incurs substantial computational overhead due to long decoding from short prompts. While sparse attention can reduce latency and memory usage, existing methods often degrade reasoning accuracy because selection err…