2018
The Mirage of Action-Dependent Baselines in Reinforcement Learning
ICML 2018oral
Policy gradient methods are a widely used class of model-free reinforcement learning algorithms where a state-dependent baseline is used to reduce gradient estimator variance. Several recent papers extend the baseline to depend on both the state and action and suggest that this significantly reduces…