Towards A Better Initial Policy Model For Scalable Long-CoT Reinforcement Learning
Long-CoT reasoning combined with reinforcement learning for large language models demonstrates remarkable performance and scalability. However, we observe that the initial policy model could significantly influence the final performance as well as the token efficiency. Additionally, there is a lack…