2020
A Hybrid Stochastic Policy Gradient Algorithm for Reinforcement Learning
AISTATS 2020poster
We propose a novel hybrid stochastic policy gradient estimator by combining an unbiased policy gradient estimator, the REINFORCE estimator, with another biased one, an adapted SARAH estimator for policy optimization. The hybrid policy gradient estimator is shown to be biased, but has variance reduce…