2026
Revisiting the Data Sampling in Multimodal Post-training from a Difficulty-Distinguish View
AAAI 2026technical
Recent advances in Multimodal Large Language Models (MLLMs) have spurred significant progress in Chain-of-Thought (CoT) reasoning. Building on the success of Deepseek-R1, researchers extended multimodal reasoning to post-training paradigms based on reinforcement learning (RL), focusing predominantly