Bridging Scaling Laws to On-Policy Reinforcement Learning via Adaptive Batch Scaling
While the "Scaling Laws" have driven massive success in Computer Vision and NLP through large-scale training with massive batch sizes, Reinforcement Learning (RL) has largely failed to benefit from this paradigm. In RL, increasing batch sizes beyond a modest threshold often leads to diminishing retu…