← Search

Alexa R. Tartaglini

2 accepted papers

2026

Addressing divergent representations from causal interventions on neural networks

ICLR 2026oral

A common approach to mechanistic interpretability is to causally manipulate model representations via targeted interventions in order to understand what those representations encode. Here we ask whether such interventions create out-of-distribution (divergent) representations, and whether this raise…

Cited by 0SourceScholar
2024

Beyond the Doors of Perception: Vision Transformers Represent Relations Between Objects

NeurIPS 2024poster

Though vision transformers (ViTs) have achieved state-of-the-art performance in a variety of settings, they exhibit surprising failures when performing tasks involving visual relations. This begs the question: how do ViTs attempt to perform tasks that require computing visual relations between objec…