2022
Model-based Safe Deep Reinforcement Learning via a Constrained Proximal Policy Optimization Algorithm
NeurIPS 2022accept
During initial iterations of training in most Reinforcement Learning (RL) algorithms, agents perform a significant number of random exploratory steps. In the real world, this can limit the practicality of these algorithms as it can lead to potentially dangerous behavior. Hence safe exploration is a…