← Search

Shiqing Gao

3 accepted papers

2025

Controlling Underestimation Bias in Constrained Reinforcement Learning for Safe Exploration

ICML 2025oral

Constrained Reinforcement Learning (CRL) aims to maximize cumulative rewards while satisfying constraints. However, existing CRL algorithms often encounter significant constraint violations during training, limiting their applicability in safety-critical scenarios. In this paper, we identify the und…

Cited by 0SourcePDFScholar
2025

Extreme Value Policy Optimization for Safe Reinforcement Learning

ICML 2025poster

Ensuring safety is a critical challenge in applying Reinforcement Learning (RL) to real-world scenarios. Constrained Reinforcement Learning (CRL) addresses this by maximizing returns under predefined constraints, typically formulated as the expected cumulative cost. However, expectation-based constr…

Cited by 0SourcePDFScholar
2024

Exterior Penalty Policy Optimization with Penalty Metric Network under Constraints

IJCAI 2024poster

In Constrained Reinforcement Learning (CRL), agents explore the environment to learn the optimal policy while satisfying constraints. The penalty function method has recently been studied as an effective approach for handling constraints, which imposes constraints penalties on the objective to trans…