2025
MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention
NeurIPS 2025spotlight
Transformers have achieved state-of-the-art performance across various tasks, but suffer from a notable quadratic complexity in sequence length due to the attention mechanism. In this work, we propose MonarchAttention -- a novel approach to sub-quadratic attention approximation via Monarch matrices,…