2025
EvolvedGRPO: Unlocking Reasoning in LVLMs via Progressive Instruction Evolution
NeurIPS 2025poster
Recent advances in reinforcement learning (RL) methods such as Grouped Relative Policy Optimization (GRPO) have strengthened the reasoning capabilities of Large Vision-Language Models (LVLMs). However, due to the inherent entanglement between visual and textual modalities, applying GRPO to LVLMs oft…