2025
Features that Make a Difference: Leveraging Gradients for Improved Dictionary Learning
NAACL 2025findings
Sparse Autoencoders (SAEs) are a promising approach for extracting neural network representations by learning a sparse and overcomplete decomposition of the network’s internal activations. However, SAEs are traditionally trained considering only activation values and not the effect those activations…