2026
Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models
ICML 2026poster
The quadratic computational complexity of softmax transformers has become a bottleneck in long-context scenarios. In contrast, linear attention model families provide a promising direction towards a more efficient sequential model. These linear attention models compress past $KV$ values into a singl…