← Search

Omar Elmansouri

1 accepted papers

2026

Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients

ICML 2026poster

Reinforcement learning from human feedback (RLHF) or verifiable rewards (RLVR), the standard paradigm for aligning LLMs or building recent SOTA reasoning models, is highly sensitive to noise from inconsistent or erroneous rewards. Yet, the interaction between such noise and widely used group-based p…

Cited by 0SourceScholar