← Search

Simon Jerome Han

1 accepted papers

2026

Addressing divergent representations from causal interventions on neural networks

ICLR 2026oral

A common approach to mechanistic interpretability is to causally manipulate model representations via targeted interventions in order to understand what those representations encode. Here we ask whether such interventions create out-of-distribution (divergent) representations, and whether this raise…

Cited by 0SourceScholar