NeurIPS 2019poster171 citations

Provably Global Convergence of Actor-Critic: A Case for Linear Quadratic Regulator with Ergodic Cost

Zhuoran Yang, Yongxin Chen, Mingyi Hong, Zhaoran Wang

Abstract

Despite the empirical success of the actor-critic algorithm, its theoretical understanding lags behind. In a broader context, actor-critic can be viewed as an online alternating update algorithm for bilevel optimization, whose convergence is known to be fragile. To understand the instability of actor-critic, we focus on its application to linear quadratic regulators, a simple yet fundamental setting of reinforcement learning. We establish a nonasymptotic convergence analysis of actor- critic in this setting. In particular, we prove that actor-critic finds a globally optimal pair of actor (policy) and critic (action-value function) at a linear rate of convergence. Our analysis may serve as a preliminary step towards a complete theoretical understanding of bilevel optimization with nonconvex subproblems, which is NP-hard in the worst case and is often solved using heuristics.

BibTeX
@inproceedings{NEURIPS2019_9713faa2,
 author = {Yang, Zhuoran and Chen, Yongxin and Hong, Mingyi and Wang, Zhaoran},
 booktitle = {Advances in Neural Information Processing Systems},
 editor = {H. Wallach and H. Larochelle and A. Beygelzimer and F. d\textquotesingle Alch\'{e}-Buc and E. Fox and R. Garnett},
 pages = {},
 publisher = {Curran Associates, Inc.},
 title = {Provably Global Convergence of Actor-Critic: A Case for Linear Quadratic Regulator with Ergodic Cost},
 url = {https://proceedings.neurips.cc/paper_files/paper/2019/file/9713faa264b94e2bf346a1bb52587fd8-Paper.pdf},
 volume = {32},
 year = {2019}
}
Provably Global Convergence of Actor-Critic: A Case for Linear Quadratic Regulator with Ergodic Cost · NeurIPS 2019