2026
Does “Do Differentiable Simulators Give Better Policy Gradients?” Give Better Policy Gradients?
ICLR 2026poster
In policy gradient reinforcement learning, access to a differentiable model enables 1st-order gradient estimation that accelerates learning compared to relying solely on derivative-free 0th-order estimators. However, discontinuous dynamics cause bias and undermine the effectiveness of 1st-order esti…