← Search

Yuhao Ding

10 accepted papers

2024

Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation

AAAI 2024technical

Ensuring the safety of Reinforcement Learning (RL) is crucial for its deployment in real-world applications. Nevertheless, managing the trade-off between reward and safety during exploration presents a significant challenge. Improving reward performance through policy adjustments may adversely affec…

2024

Enhancing Efficiency of Safe Reinforcement Learning via Sample Manipulation

NeurIPS 2024poster

Safe reinforcement learning (RL) is crucial for deploying RL agents in real-world applications, as it aims to maximize long-term rewards while satisfying safety constraints. However, safe RL often suffers from sample inefficiency, requiring extensive interactions with the environment to learn a safe…

2023

A CMDP-within-online framework for Meta-Safe Reinforcement Learning

ICLR 2023top-25%

Meta-reinforcement learning has widely been used as a learning-to-learn framework to solve unseen tasks with limited experience. However, the aspect of constraint violations has not been adequately addressed in the existing works, making their application restricted in real-world settings. In this p…

Cited by 22SourcePDFScholar
2023

Non-stationary Risk-Sensitive Reinforcement Learning: Near-Optimal Dynamic Regret, Adaptive Detection, and Separation Design

AAAI 2023technical

We study risk-sensitive reinforcement learning (RL) based on an entropic risk measure in episodic non-stationary Markov decision processes (MDPs). Both the reward functions and the state transition kernels are unknown and allowed to vary arbitrarily over time with a budget on their cumulative variat…

Cited by 8SourcePDFScholar
2023

Policy-Based Primal-Dual Methods for Convex Constrained Markov Decision Processes

AAAI 2023technical

We study convex Constrained Markov Decision Processes (CMDPs) in which the objective is concave and the constraints are convex in the state-action occupancy measure. We propose a policy-based primal-dual algorithm that updates the primal variable via policy gradient ascent and updates the dual varia…

Cited by 15SourcePDFScholar
2023

Provably Efficient Primal-Dual Reinforcement Learning for CMDPs with Non-stationary Objectives and Constraints

AAAI 2023technical

We consider primal-dual-based reinforcement learning (RL) in episodic constrained Markov decision processes (CMDPs) with non-stationary objectives and constraints, which plays a central role in ensuring the safety of RL in time-varying environments. In this problem, the reward/utility functions and…

Cited by 35SourcePDFScholar
2023

Scalable Primal-Dual Actor-Critic Method for Safe Multi-Agent RL with General Utilities

NeurIPS 2023poster

We investigate safe multi-agent reinforcement learning, where agents seek to collectively maximize an aggregate sum of local objectives while satisfying their own safety constraints. The objective and constraints are described by general utilities, i.e., nonlinear functions of the long-term state-ac…

Cited by 16SourcePDFScholar
2023

Tempo Adaptation in Non-stationary Reinforcement Learning

NeurIPS 2023poster

We first raise and tackle a ``time synchronization'' issue between the agent and the environment in non-stationary reinforcement learning (RL), a crucial factor hindering its real-world applications. In reality, environmental changes occur over wall-clock time ($t$) rather than episode progress ($k$…

2022

A Dual Approach to Constrained Markov Decision Processes with Entropy Regularization

AISTATS 2022poster

We study entropy-regularized constrained Markov decision processes (CMDPs) under the soft-max parameterization, in which an agent aims to maximize the entropy-regularized value function while satisfying constraints on the expected total utility. By leveraging the entropy regularization, our theoreti…

Cited by 44SourcePDFScholar