2026
Cumulant Attention in Vision Transformers (Student Abstract)
AAAI 2026technical
Transformer models have achieved remarkable success across diverse deep learning fields, including natural language processing (NLP) and computer vision (CV). One drawback of these models is that the computational cost of the softmax attention, the core component of the transformer, exhibits quadrat