2025
Integral Transformer: Denoising Attention, Not Too Much Not Too Little
EMNLP 2025
Softmax self-attention often assigns disproportionate weight to semantically uninformative tokens such as punctuation and special tokens, a phenomenon known as attention noise. While recent methods like Cog Attention and the Differential Transformer have addressed this by introducing negative attent