← Search

Constantin Venhoff

5 accepted papers

2026

Base Models Know How to Reason, Thinking Models Learn When

ICML 2026spotlight

Why do thinking language models outperform their base counterparts, and what exactly do they learn during training? We introduce constructive model diffing, a framework for understanding fine-tuned models by explicitly constructing the base-to-fine-tuned difference from interpretable components to p…

Cited by 0SourceScholar
2026

Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP

ICLR 2026poster

Typographic attacks exploit multi-modal systems by injecting text into images, leading to targeted misclassifications, malicious content generation and even Vision-Language Model jailbreaks. In this work, we analyze how CLIP vision encoders behave under typographic attacks, locating specialized atte…

Cited by 0SourcecodeScholar
2026

Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive Learning

ICLR 2026poster

Contrastive Language-Image Pre-training (CLIP) has become a cornerstone in vision-language representation learning, powering diverse downstream tasks and serving as the default vision backbone in multimodal large language models (MLLMs). Despite its success, CLIP's dense and opaque latent representa…

Cited by 0SourceScholar
2025

Mixture of Experts Made Intrinsically Interpretable

ICML 2025poster

Neurons in large language models often exhibit \emph{polysemanticity}, simultaneously encoding multiple unrelated concepts and obscuring interpretability. Instead of relying on post-hoc methods, we present \textbf{MoE-X}, a mixture-of-experts (MoE) language model designed to be \emph{intrinsically}…

Cited by 0SourcePDFScholar
2025

Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval

NeurIPS 2025poster

Training vision language models (VLMs) aims to align visual representations from a vision encoder with the textual representations of a pretrained large language model (LLM). However, many VLMs exhibit reduced factual recall performance compared to their LLM backbones, raising the question of how ef…

Cited by 0SourceScholar