2026
D²Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement Learning
ICML 2026poster
Reinforcement learning (RL) has demonstrated potential for enhancing reasoning in large language models (LLMs). However, effective RL training, which requires medium-difficulty training samples, faces two fundamental challenges: Effective Data Scarcity and Dynamic Difficulty Shifts, where medium-dif…