← Search

Yuzhang Luo

1 accepted papers

2026

Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units

ICML 2026oral

Mechanistic Interpretability has successfully identified functional circuits in Large Language Models (LLMs), yet their causal origins in the training data remain poorly understood. We bridge this gap by introducing **Mechanistic Data Attribution (MDA)**, a scalable framework that traces the formati…

Cited by 0SourceScholar