← Search

Marten Dijk

2 accepted papers

2020

A Hybrid Stochastic Policy Gradient Algorithm for Reinforcement Learning

AISTATS 2020poster

We propose a novel hybrid stochastic policy gradient estimator by combining an unbiased policy gradient estimator, the REINFORCE estimator, with another biased one, an adapted SARAH estimator for policy optimization. The hybrid policy gradient estimator is shown to be biased, but has variance reduce…

2018

SGD and Hogwild! Convergence Without the Bounded Gradients Assumption

ICML 2018oral

Stochastic gradient descent (SGD) is the optimization algorithm of choice in many machine learning applications such as regularized empirical risk minimization and training deep neural networks. The classical convergence analysis of SGD is carried out under the assumption that the norm of the stocha…

Cited by 266SourcePDFScholar