Correlations in the Data Lead to Semantically Rich Feature Geometry Under Superposition
Recent advances in mechanistic interpretability have shown that many features represented by deep learning models can be captured by dictionary learning approaches such as sparse autoencoders. However, our understanding of the structures formed by these internal representations is still limited. Ini…