← Search

Bruno Puri

2 accepted papers

2026

Circuit Insights: Towards Interpretability Beyond Activations

ICLR 2026poster

The fields of explainable AI and mechanistic interpretability aim to uncover the internal structure of neural networks, with circuit discovery as a central tool for understanding model computations. Existing approaches, however, rely on manual inspection and remain limited to toy tasks. Automated in…

Cited by 0SourceScholar
2025

FADE: Why Bad Descriptions Happen to Good Features

ACL 2025finding

Recent advances in mechanistic interpretability have highlighted the potential of automating interpretability pipelines in analyzing the latent representations within LLMs. While this may enhance our understanding of internal mechanisms, the field lacks standardized evaluation methods for assessing…