NeurIPS 2018poster115 citations

Constrained Cross-Entropy Method for Safe Reinforcement Learning

Min Wen, Ufuk Topcu

Abstract

We study a safe reinforcement learning problem in which the constraints are defined as the expected cost over finite-length trajectories. We propose a constrained cross-entropy-based method to solve this problem. The method explicitly tracks its performance with respect to constraint satisfaction and thus is well-suited for safety-critical applications. We show that the asymptotic behavior of the proposed algorithm can be almost-surely described by that of an ordinary differential equation. Then we give sufficient conditions on the properties of this differential equation to guarantee the convergence of the proposed algorithm. At last, we show with simulation experiments that the proposed algorithm can effectively learn feasible policies without assumptions on the feasibility of initial policies, even with non-Markovian objective functions and constraint functions.

BibTeX
@inproceedings{NEURIPS2018_34ffeb35,
 author = {Wen, Min and Topcu, Ufuk},
 booktitle = {Advances in Neural Information Processing Systems},
 editor = {S. Bengio and H. Wallach and H. Larochelle and K. Grauman and N. Cesa-Bianchi and R. Garnett},
 pages = {},
 publisher = {Curran Associates, Inc.},
 title = {Constrained Cross-Entropy Method for Safe Reinforcement Learning},
 url = {https://proceedings.neurips.cc/paper_files/paper/2018/file/34ffeb359a192eb8174b6854643cc046-Paper.pdf},
 volume = {31},
 year = {2018}
}
Constrained Cross-Entropy Method for Safe Reinforcement Learning · NeurIPS 2018