2026
A Provable Expressiveness Hierarchy in Hybrid Linear-Full Attention
ICML 2026poster
Transformers serve as the foundation of most modern large language models. To mitigate the quadratic complexity of standard full attention, various efficient attention mechanisms, such as linear and hybrid attention, have been developed. A fundamental gap remains: their expressive power relative to …