2025
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
ICML 2025poster
Sparse autoencoders (SAEs) have emerged as a powerful tool for interpreting neural networks by extracting the concepts represented in their activations. However, choosing the size of the SAE dictionary (i.e. number of learned concepts) creates a tension: as dictionary size increases to capture more…