← Search

Dana Arad

5 accepted papers

2025

MIB: A Mechanistic Interpretability Benchmark

ICML 2025poster

How can we know whether new mechanistic interpretability methods achieve real improvements? In pursuit of lasting evaluation standards, we propose MIB, a Mechanistic Interpretability Benchmark, with two tracks spanning four tasks and five models. MIB favors methods that precisely and concisely recov…

2025

Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs

NeurIPS 2025poster

Vision-Language models (VLMs) show impressive abilities to answer questions on visual inputs (e.g., counting objects in an image), yet demonstrate higher accuracies when performing an analogous task on text (e.g., counting words in a text). We investigate this accuracy gap by identifying and compari…

Cited by 0SourcecodeScholar
2024

Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines

ACL 2024long

Text-to-image diffusion models (T2I) use a latent representation of a text prompt to guide the image generation process. However, the process by which the encoder produces the text representation is unknown. We propose the Diffusion Lens, a method for analyzing the text encoder of T2I models by gene…

Cited by 8SourcePDFScholar
2024

ReFACT: Updating Text-to-Image Models by Editing the Text Encoder

NAACL 2024long

Our world is marked by unprecedented technological, global, and socio-political transformations, posing a significant challenge to textto-image generative models. These models encode factual associations within their parameters that can quickly become outdated, diminishing their utility for end-user…