Dynamics-Predictive Sampling for Active RL Finetuning of Large Reasoning Models
Reinforcement learning (RL) finetuning has become a key technique for enhancing the reasoning abilities of large language models (LLMs). However, its effectiveness critically depends on the selection of training data. Recent advances underscore the importance of online prompt selection methods, whic…