ICLR 2026poster0 citations

Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning

Shubham Parashar, Shurui Gui, Xiner Li, Hongyi Ling, Sushil Vemuri, Blake Olson, Eric Li, Yu Zhang

Abstract

We aim to improve the reasoning capabilities of language models via reinforcement learning with verifiable rewards (RLVR). Recent RLVR post-trained models like DeepSeek-R1 have demonstrated reasoning abilities on mathematical and coding tasks. However, prior studies suggest that using RLVR alone to improve reasoning on inherently difficult tasks is less effective due to sparse rewards. Here, we draw inspiration from curriculum learning and propose to schedule tasks from easy to hard (E2H), allowing LLMs to build reasoning skills gradually. Our method is termed E2H Reasoner. Empirically, we observe that, although easy tasks are important initially, fading them out through appropriate scheduling is essential in preventing overfitting. Theoretically, we establish convergence guarantees for E2H Reasoner within an approximate policy iteration framework. We derive finite-sample complexity bounds and show that when tasks are appropriately decomposed and conditioned, learning through curriculum stages requires fewer total samples than direct learning. Experiments across diverse datasets and models demonstrate that E2H Reasoner substantially enhances LLM reasoning.

LLMReinforcement LearningPost Training
BibTeX
@inproceedings{
parashar2026curriculum,
title={Curriculum Reinforcement Learning from Easy to Hard Tasks Improves {LLM} Reasoning},
author={Shubham Parashar and Shurui Gui and Xiner Li and Hongyi Ling and Sushil Vemuri and Blake Olson and Eric Li and Yu Zhang and James Caverlee and Dileep Kalathil and Shuiwang Ji},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=KJvHnl3kUv}
}
Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning · ICLR 2026