← Search

Chenxiao Zhao

2 accepted papers

2026

DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

ICLR 2026poster

Large Vision-Language Models excel at multimodal understanding but struggle to deeply integrate visual information into their predominantly text-based reasoning processes, a key challenge in mirroring human cognition. To address this, we introduce DeepEyes, a model that learns to ``think with images…

Cited by 0SourcecodeScholar
2026

DeepEyesV2: Toward Agentic Multimodal Model

ICLR 2026poster

Agentic multimodal models should not only comprehend text and images, but also actively invoke external tools, such as code execution environments and web search, and integrate these operations into reasoning. In this work, we introduce DeepEyesV2 and explore how to build an agentic multimodal model…

Cited by 0SourcecodeScholar