2026
Trust-Region Diffusion Policies for Massively Parallel On-Policy RL
ICML 2026poster
Reinforcement learning with massively parallel simulations has become an emerging trend; however, most existing approaches still rely on simple Gaussian policy parameterizations. Diffusion models provide a more expressive policy class and have shown strong performance on challenging control problems…