2025
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection
NeurIPS 2025poster
A primary impediment to scaling reinforcement learning (RL) for large language model (LLM) training is the substantial computational cost, predominantly arising from the necessity of multi-sampling for policy optimization and evaluation. This underscores the critical yet challenging nature of effici…