2026
Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention
ICML 2026poster
Transformer attention is typically implemented using softmax normalization, which enforces attention weights with unit sum normalization. While effective in many settings, this constraint can limit flexibility in controlling attention magnitudes and may contribute to overly concentrated or unstable …