2025
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
ICLR 2025poster
As AI systems are increasingly deployed in high-stakes applications, ensuring their interpretability is essential. Mechanistic Interpretability (MI) aims to reverse-engineer neural networks by extracting human-understandable algorithms embedded within their structures to explain their behavior. This…