Anatomy of Massive Activations and Attention Sinks
We study two recurring phenomena in Transformer language models. First, \emph{massive activations}, where a small number of hidden channels attain extremely large values for a few tokens. Second, \emph{attention sinks}, where certain tokens attract a disproportionate share of attention across many h…