2026
Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
ICML 2026oral
Mechanistic Interpretability has successfully identified functional circuits in Large Language Models (LLMs), yet their causal origins in the training data remain poorly understood. We bridge this gap by introducing **Mechanistic Data Attribution (MDA)**, a scalable framework that traces the formati…