← Search

Haoming Lyu

1 accepted papers

2026

SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward

ICLR 2026poster

Recent advances have shown success in eliciting strong reasoning abilities in multimodal large language models (MLLMs) through rule-based reinforcement learning (RL) with outcome rewards. However, this paradigm typically lacks supervision over the thinking process leading to the final outcome. As a…

Cited by 0SourcecodeScholar