2026
XRPO: Pushing the Limits of GRPO with Targeted Exploration and Exploitation
ICML 2026poster
Reinforcement learning algorithms such as GRPO have driven recent advances in large language model (LLM) reasoning. While scaling the number of rollouts stabilizes training, existing approaches suffer from limited exploration on challenging prompts and leave informative feedback signals underexploit…