2025
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
ACL 2025long
Long-context modeling is crucial for next-generation language models, yet the high computational cost of standard attention mechanisms poses significant computational challenges. Sparse attention offers a promising direction for improving efficiency while maintaining model capabilities. We present N…