2024
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
NeurIPS 2024poster
What latent features are encoded in language model (LM) representations? Recent work on training sparse autoencoders (SAEs) to disentangle interpretable features in LM representations has shown significant promise. However, evaluating the quality of these SAEs is difficult because we lack a ground-t…