2019
Trajectory-Based Off-Policy Deep Reinforcement Learning
ICML 2019oral
Policy gradient methods are powerful reinforcement learning algorithms and have been demonstrated to solve many complex tasks. However, these methods are also data-inefficient, afflicted with high variance gradient estimates, and frequently get stuck in local optima. This work addresses these weakne…