← Search

Lijing Shao

1 accepted papers

2026

Probing RLVR Training Instability through the Lens of Objective-Level Hacking

ICML 2026poster

Prolonged reinforcement learning with verifiable rewards (RLVR) has been shown to drive continuous improvements in the reasoning capabilities of large language models, but the training is often prone to instabilities, especially in Mixture-of-Experts (MoE) architectures. Training instability severel…

Cited by 0SourceScholar