2025
Staggered Environment Resets Improve Massively Parallel On-Policy Reinforcement Learning
NeurIPS 2025poster
Massively parallel GPU simulation environments have accelerated reinforcement learning (RL) research by enabling fast data collection for on-policy RL algorithms like Proximal Policy Optimization (PPO). To maximize throughput, it is common to use short rollouts per policy update, increasing the upda…