2025
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
ICLR 2025poster
Disentangling model activations into human-interpretable features is a central problem in interpretability. Sparse autoencoders (SAEs) have recently attracted much attention as a scalable unsupervised approach to this problem. However, our imprecise understanding of ground-truth features in realisti…