2026
A Unified Sparse Attention via Multi-Granularity Compression
ICML 2026poster
Efficient long-context understanding is increasingly vital for large language model (LLM) applications such as multi-turn dialogue and program analysis. However, the core self-attention scales quadratically with sequence length, creating a fundamental computational bottleneck. Existing sparse attent…