2026
iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning
ICML 2026poster
While visually grounded Chain-of-Thought (CoT) has emerged as a promising paradigm to enhance fine-grained perception in multimodal large language models (MLLMs), its efficacy during the inference phase remains under-scrutinized. In this work, we empirically find that mandating the explicit object b…