2026
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
ICLR 2026poster
Enhancing the multimodal reasoning capabilities of Multimodal Large Language Models (MLLMs) is a challenging task that has attracted increasing attention in the community. Recently, several studies have applied Reinforcement Learning with Verifiable Rewards (RLVR) to the multimodal domain in order t…