Dual Policy Iteration
Wen Sun, Geoffrey J. Gordon, Byron Boots, J. Bagnell
Abstract
Recently, a novel class of Approximate Policy Iteration (API) algorithms have demonstrated impressive practical performance (e.g., ExIt from [1], AlphaGo-Zero from [2]). This new family of algorithms maintains, and alternately optimizes, two policies: a fast, reactive policy (e.g., a deep neural network) deployed at test time, and a slow, non-reactive policy (e.g., Tree Search), that can plan multiple steps ahead. The reactive policy is updated under supervision from the non-reactive policy, while the non-reactive policy is improved with guidance from the reactive policy. In this work we study this Dual Policy Iteration (DPI) strategy in an alternating optimization framework and provide a convergence analysis that extends existing API theory. We also develop a special instance of this framework which reduces the update of non-reactive policies to model-based optimal control using learned local models, and provides a theoretically sound way of unifying model-free and model-based RL approaches with unknown dynamics. We demonstrate the efficacy of our approach on various continuous control Markov Decision Processes.
BibTeX
@inproceedings{NEURIPS2018_15e122e8,
author = {Sun, Wen and Gordon, Geoffrey J and Boots, Byron and Bagnell, J.},
booktitle = {Advances in Neural Information Processing Systems},
editor = {S. Bengio and H. Wallach and H. Larochelle and K. Grauman and N. Cesa-Bianchi and R. Garnett},
pages = {},
publisher = {Curran Associates, Inc.},
title = {Dual Policy Iteration},
url = {https://proceedings.neurips.cc/paper_files/paper/2018/file/15e122e839dfdaa7ce969536f94aecf6-Paper.pdf},
volume = {31},
year = {2018}
}