How Far Can Unsupervised RLVR Scale LLM Training?
Yuxin Zuo, Bingxiang He, Zeyuan Liu, Shangziqi Zhao, Zixuan Fu, Junlin Yang, Kaiyan Zhang, Yuchen Fan
Abstract
Unsupervised Reinforcement Learning with Verifiable Rewards (URLVR) offers a pathway for Large Language Models (LLMs) to improve without human supervision. Particularly, many works use model intrinsic information as rewards for URLVR, showing promising improvements, yet their potential and limitations remain unclear. In this work, we revisit URLVR through the lens of intrinsic rewards. We present a unified theoretical framework showing that intrinsic reward methods share a core mechanism: they trade uncertainty for performance by leveraging the model’s prior knowledge to sharpen output distributions. Empirical analysis confirms this tradeoff, revealing distinct failure modes and showing that collapse is not inevitable in small, domain-specific regimes such as test-time training. Beyond these findings, early intrinsic reward dynamics also provide a lightweight indicator of model-task priors, complementing $pass@k$ in assessing RL trainability. These insights highlight both the promise and pitfalls of URLVR, motivating future directions such as external rewards and hybrid supervision strategies.
BibTeX
@inproceedings{
zuo2026how,
title={How Far Can Unsupervised {RLVR} Scale {LLM} Training?},
author={Yuxin Zuo and Bingxiang He and Zeyuan Liu and Shangziqi Zhao and Zixuan Fu and Junlin Yang and Kaiyan Zhang and Yuchen Fan and Ganqu Cui and Cheng Qian and Xiusi Chen and Youbang Sun and Xingtai Lv and Xuekai Zhu and Li Sheng and Ran Li and Huan-ang Gao and Yuchen Zhang and Lifan Yuan and Zhiyuan Liu and Bowen Zhou and Ning Ding},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=VesLZukY5E}
}