2026
Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training
ICML 2026poster
Post-training with explicit reasoning traces is common to improve the reasoning capabilities of Multimodal Large Language Models (MLLMs). However, acquiring high-quality reasoning traces is often costly and time-consuming. Hence, the self-improvement paradigm has emerged, enabling MLLMs to self-gene…