← Search

Anobel Odisho

2 accepted papers

2025

Efficient Automated Circuit Discovery in Transformers using Contextual Decomposition

ICLR 2025poster

Automated mechanistic interpretation research has attracted great interest due to its potential to scale explanations of neural network internals to large models. Existing automated circuit discovery work relies on activation patching or its approximations to identify subgraphs in models for specifi…

Cited by 1SourcePDFScholar
2024

Diagnosing Transformers: Illuminating Feature Spaces for Clinical Decision-Making

ICLR 2024poster

Pre-trained transformers are often fine-tuned to aid clinical decision-making using limited clinical notes. Model interpretability is crucial, especially in high-stakes domains like medicine, to establish trust and ensure safety, which requires human engagement. We introduce SUFO, a systematic frame…