2026
Falsifying Sparse Autoencoder Reasoning Features in Language Models
ICML 2026poster
We study how reliably sparse autoencoders (SAEs) support claims about reasoning-related internal features in large language models. We first give a stylized analysis showing that sparsity-regularized decoding can preferentially retain stable low-dimensional correlates while suppressing high-dimensio…