2026
Activating Visual Context and Commonsense Reasoning Through Masked Prediction in VLMs
AAAI 2026technical
Recent breakthroughs in reasoning models have markedly advanced the reasoning capabilities of large language models, particularly via training on tasks with verifiable rewards. Yet, a significant gap persists in their adaptation to real-world multimodal scenarios, most notably, vision-language tasks