2025
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
ICML 2025poster
We introduce a new approach to systematically map features discovered by sparse autoencoder across consecutive layers of large language models, extending earlier work that examined inter-layer feature links. By using a data-free cosine similarity technique, we trace how specific features persist, tr…