2026
On the Relationship Between Activation Outliers and Feature Death in Sparse Autoencoders
ICML 2026poster
Sparse autoencoders (SAEs) decompose neural network activations into interpretable features, but many features never activate- a problem called feature death. Death rates vary dramatically across models: near-zero on GPT-2, over 70\% on AlphaFold3 with identical SAE configurations. Why? We find that…