2022
Human-Feedback Shield Synthesis for Perceived Safety in Deep Reinforcement Learning
RA-L 2022
Despite the successes of deep reinforcement learning (RL), it is still challenging to obtain safe policies. Formal verification approaches ensure safety at all times, but usually overly restrict the agent’s behaviors, since they assume adversarial behavior of the environment. Instead of assuming adv