2022
Constrained Variational Policy Optimization for Safe Reinforcement Learning
ICML 2022spotlight
Safe reinforcement learning (RL) aims to learn policies that satisfy certain constraints before deploying them to safety-critical applications. Previous primal-dual style approaches suffer from instability issues and lack optimality guarantees. This paper overcomes the issues from the perspective of…