2024
Observable Propagation: Uncovering Feature Vectors in Transformers
ICML 2024poster
A key goal of current mechanistic interpretability research in NLP is to find *linear features* (also called "feature vectors") for transformers: directions in activation space corresponding to concepts that are used by a given model in its computation. Present state-of-the-art methods for finding l…