2018
Action-dependent Control Variates for Policy Optimization via Stein Identity
ICLR 2018poster
Policy gradient methods have achieved remarkable successes in solving challenging reinforcement learning problems. However, it still often suffers from the large variance issue on policy gradient estimation, which leads to poor sample efficiency during training. In this work, we propose a control va…