← Search

Iván Arcuschin

6 accepted papers

2026

Base Models Know How to Reason, Thinking Models Learn When

ICML 2026spotlight

Why do thinking language models outperform their base counterparts, and what exactly do they learn during training? We introduce constructive model diffing, a framework for understanding fine-tuned models by explicitly constructing the base-to-fine-tuned difference from interpretable components to p…

Cited by 0SourceScholar
2026

Biases in the Blind Spot: Detecting What LLMs Fail to Mention

ICML 2026poster

Large Language Models (LLMs) often provide chain-of-thought (CoT) reasoning traces that appear plausible, but may hide internal biases. We call these *unverbalized biases*. Monitoring models via their stated reasoning is therefore unreliable, and existing bias evaluations typically require predefine…

Cited by 0SourceScholar
2026

Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

ICML 2026poster

Recent studies indicate that when faced with explicit biases in prompts, models often omit mentioning these biases in their Chain-of-Thought (CoT) output, revealing that verbalized reasoning can give an incorrect picture of how models arrive at conclusions (unfaithfulness). In this work, we show tha…

Cited by 0SourcecodeScholar
2025

MIB: A Mechanistic Interpretability Benchmark

ICML 2025poster

How can we know whether new mechanistic interpretability methods achieve real improvements? In pursuit of lasting evaluation standards, we propose MIB, a Mechanistic Interpretability Benchmark, with two tracks spanning four tasks and five models. MIB favors methods that precisely and concisely recov…

2024

InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques

NeurIPS 2024poster

Mechanistic interpretability methods aim to identify the algorithm a neural network implements, but it is difficult to validate such methods when the true algorithm is unknown. This work presents InterpBench, a collection of semi-synthetic yet realistic transformers with known circuits for evaluatin…

Cited by 4SourcePDFScholar