2026
DSSA: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
ICLR 2026poster
Long-sequence processing is a critical capability for modern large language models. However, the self-attention mechanism in the standard Transformer architecture faces severe computational and memory bottlenecks when processing long sequences. While trainable sparse attention methods offer a promis…