2025
Validating Mechanistic Interpretations: An Axiomatic Approach
ICML 2025poster
Mechanistic interpretability aims to reverse engineer the computation performed by a neural network in terms of its internal components. Although there is a growing body of research on mechanistic interpretation of neural networks, the notion of a *mechanistic interpretation* itself is often ad-hoc.…