2026
Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
ICML 2026poster
Reinforcement learning from human feedback (RLHF) or verifiable rewards (RLVR), the standard paradigm for aligning LLMs or building recent SOTA reasoning models, is highly sensitive to noise from inconsistent or erroneous rewards. Yet, the interaction between such noise and widely used group-based p…