2025
Variance Sensitivity Induces Attention Entropy Collapse and Instability in Transformers
EMNLP 2025
Attention-based language models commonly rely on the softmax function to convert attention logits into probability distributions. However, this softmax re-weighting can lead to *attention entropy collapse*, in which attention disproportionately concentrates on a single token, ultimately causing trai