2024
Sparse Autoencoders Find Highly Interpretable Features in Language Models
ICLR 2024poster
One of the roadblocks to a better understanding of neural networks' internals is \textit{polysemanticity}, where neurons appear to activate in multiple, semantically distinct contexts. Polysemanticity prevents us from identifying concise, human-understandable explanations for what neural networks ar…