← Search

Sumanta Dey

1 accepted papers

2024

P2BPO: Permeable Penalty Barrier-Based Policy Optimization for Safe RL

AAAI 2024technical

Safe Reinforcement Learning (SRL) algorithms aim to learn a policy that maximizes the reward while satisfying the safety constraints. One of the challenges in SRL is that it is often difficult to balance the two objectives of reward maximization and safety constraint satisfaction. Existing algorithm…