2026
Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
ICLR 2026poster
Vision-language reinforcement learning (RL) has primarily focused on narrow domains (e.g. geometry or chart reasoning). This leaves broader training scenarios and resources underexplored, limiting the exploration and learning of Vision Language Models (VLMs) through RL. We find video games inherentl…