2026
VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
ICLR 2026poster
Reinforcement learning finetuning (RFT) has significantly advanced the reasoning capabilities of large language models (LLMs) by enabling long chains of thought, multi-turn self-correction, and effective tool use. While recent works attempt to extend RFT to vision-language models (VLMs), these effor…