2026
See First, Reason Later: Mutual Information-Guided Reinforcement Learning for Vision-Language Models
ICML 2026poster
Vision-Language Models (VLMs) frequently suffer from visual perception errors and hallucinations that compromise answer accuracy in complex reasoning tasks. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising solution by optimizing policies using answer correctness signals. Desp…