2026
Singular Vectors of Attention Heads Align with Features
ICML 2026poster
Identifying feature representations in language models is a central task in mechanistic interpretability. Several recent studies have made an implicit assumption that feature representations can be inferred in some cases from singular vectors of attention matrices. However, sound justification for t…