ICLR 2023poster51 citations

Linear Convergence of Natural Policy Gradient Methods with Log-Linear Policies

Rui Yuan, Simon Shaolei Du, Robert M. Gower, Alessandro Lazaric, Lin Xiao

Abstract

We consider infinite-horizon discounted Markov decision processes and study the convergence rates of the natural policy gradient (NPG) and the Q-NPG methods with the log-linear policy class. Using the compatible function approximation framework, both methods with log-linear policies can be written as approximate versions of the policy mirror descent (PMD) method. We show that both methods attain linear convergence rates and $\tilde{\mathcal{O}}(1/\epsilon^2)$ sample complexities using a simple, non-adaptive geometrically increasing step size, without resorting to entropy or other strongly convex regularization. Lastly, as a byproduct, we obtain sublinear convergence rates for both methods with arbitrary constant step size.

Discounted Markov decision processnatural policy gradientpolicy mirror descentlog-linear policysample complexity
BibTeX
@inproceedings{
yuan2023linear,
title={Linear Convergence of Natural Policy Gradient Methods with Log-Linear Policies},
author={Rui Yuan and Simon Shaolei Du and Robert M. Gower and Alessandro Lazaric and Lin Xiao},
booktitle={The Eleventh International Conference on Learning Representations },
year={2023},
url={https://openreview.net/forum?id=-z9hdsyUwVQ}
}
Linear Convergence of Natural Policy Gradient Methods with Log-Linear Policies · ICLR 2023