2025
Enhancing Efficiency and Exploration in Reinforcement Learning for LLMs
EMNLP 2025
Reasoning large language models (LLMs) excel in complex tasks, which has drawn significant attention to reinforcement learning (RL) for LLMs. However, existing approaches allocate an equal number of rollouts to all questions during the RL process, which is inefficient. This inefficiency stems from t