2026
CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization
CVPR 2026
Agentic vision-language models are increasingly trained to "think with images" by calling image operations. However, we show that high final-answer accuracy often hides unfaithful visual reasoning: models may invoke tools on irrelevant regions or ignore tool outputs entirely, yet still guess the cor