2025
Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment
ICML 2025poster
We present Universal Sparse Autoencoders (USAEs), a framework for uncovering and aligning interpretable concepts spanning multiple pretrained deep neural networks. Unlike existing concept-based interpretability methods, which focus on a single model, USAEs jointly learn a universal concept space tha…