2022
DOPE: Doubly Optimistic and Pessimistic Exploration for Safe Reinforcement Learning
NeurIPS 2022accept
Safe reinforcement learning is extremely challenging--not only must the agent explore an unknown environment, it must do so while ensuring no safety constraint violations. We formulate this safe reinforcement learning (RL) problem using the framework of a finite-horizon Constrained Markov Decision…