← Search

Georg Lange

2 accepted papers

2025

Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control

ICLR 2025poster

Disentangling model activations into human-interpretable features is a central problem in interpretability. Sparse autoencoders (SAEs) have recently attracted much attention as a scalable unsupervised approach to this problem. However, our imprecise understanding of ground-truth features in realisti…

Cited by 30SourcePDFScholar
2024

Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching

ICLR 2024poster

Mechanistic interpretability aims to attribute high-level model behaviors to specific, interpretable learned features. It is hypothesized that these features manifest as directions or low-dimensional subspaces within activation space. Accordingly, recent studies have explored the identification and…