2026
Probing RLVR Training Instability through the Lens of Objective-Level Hacking
ICML 2026poster
Prolonged reinforcement learning with verifiable rewards (RLVR) has been shown to drive continuous improvements in the reasoning capabilities of large language models, but the training is often prone to instabilities, especially in Mixture-of-Experts (MoE) architectures. Training instability severel…