2025
Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders
ICML 2025poster
A recent line of work has shown promise in using sparse autoencoders (SAEs) to uncover interpretable features in neural network representations. However, the simple linear-nonlinear encoding mechanism in SAEs limits their ability to perform accurate sparse inference. Using compressed sensing theory,…