2025
Scaling Sparse Feature Circuits For Studying In-Context Learning
ICML 2025poster
Sparse autoencoders (SAEs) are a popular tool for interpreting large language model activations, but their utility in addressing open questions in interpretability remains unclear. In this work, we demonstrate their effectiveness by using SAEs to deepen our understanding of the mechanism behind in-c…