2025
The Logical Implication Steering Method for Conditional Interventions on Transformer Generation
ICML 2025poster
The field of mechanistic interpretability in pre-trained transformer models has demonstrated substantial evidence supporting the ''linear representation hypothesis'', which is the idea that high level concepts are encoded as vectors in the space of activations of a model. Studies also show that mode…