2026
What You Think is What You See: Driving Exploration in VLM Agents via Visual-Linguistic Curiosity
ICML 2026spotlight
To navigate partially observable visual environments, recent VLM agents increasingly internalize world modeling capabilities directly into their policies via explicit CoT reasoning with reinforcement learning (RL). However, mere passive exploitation of reasoning on visited states is insufficient for…