2026
Taming Polysemanticity in LLMs: Theory-Grounded Feature Recovery via Sparse Autoencoders
ICLR 2026poster
We study the challenge of achieving theoretically grounded feature recovery using Sparse Autoencoders (SAEs) for the interpretation of Large Language Models. Existing SAE training algorithms often lack rigorous mathematical guarantees and suffer from practical limitations such as hyperparameter sen…