2026
SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward
ICLR 2026poster
Recent advances have shown success in eliciting strong reasoning abilities in multimodal large language models (MLLMs) through rule-based reinforcement learning (RL) with outcome rewards. However, this paradigm typically lacks supervision over the thinking process leading to the final outcome. As a…