2026
Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining
ICML 2026poster
We present Multipole Semantic Attention (MuSe), an efficient approximation of softmax attention for long-context transformers. MuSe clusters queries and keys separately in their learned representation spaces, computing query-specific cluster summaries that capture how each query cluster attends to e…