← Search

Archana Bura

2 accepted papers

2022

DOPE: Doubly Optimistic and Pessimistic Exploration for Safe Reinforcement Learning

NeurIPS 2022accept

Safe reinforcement learning is extremely challenging--not only must the agent explore an unknown environment, it must do so while ensuring no safety constraint violations. We formulate this safe reinforcement learning (RL) problem using the framework of a finite-horizon Constrained Markov Decision…

2021

Learning with Safety Constraints: Sample Complexity of Reinforcement Learning for Constrained MDPs

AAAI 2021technical

Many physical systems have underlying safety considerations that require that the policy employed ensures the satisfaction of a set of constraints. The analytical formulation usually takes the form of a Constrained Markov Decision Process (CMDP). We focus on the case where the CMDP is unknown, and…

Cited by 52SourcePDFScholar