← Search

Kevin Ro Wang

1 accepted papers

2023

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small

ICLR 2023poster

Research in mechanistic interpretability seeks to explain behaviors of ML models in terms of their internal components. However, most previous work either focuses on simple behaviors in small models, or describes complicated behaviors in larger models with broad strokes. In this work, we bridge this…