2020
Sample Efficient Policy Gradient Methods with Recursive Variance Reduction
ICLR 2020poster
Improving the sample efficiency in reinforcement learning has been a long-standing research problem. In this work, we aim to reduce the sample complexity of existing policy gradient methods. We propose a novel policy gradient algorithm called SRVR-PG, which only requires $O(1/\epsilon^{3/2})$\footno…