2026
Thought Branches: Interpreting LLM Reasoning Requires Resampling
ICLR 2026poster
We argue that interpreting reasoning models from a single chain-of-thought (CoT) is fundamentally inadequate. To understand computation and causal influence, one must study reasoning as a distribution of possible trajectories elicited by a given prompt. We approximate this distribution via on-policy…