← Search

Donghao Ying

6 accepted papers

2025

Don’t Trade Off Safety: Diffusion Regularization for Constrained Offline RL

NeurIPS 2025poster

Constrained reinforcement learning (RL) seeks high-performance policies under safety constraints. We focus on an offline setting where the agent learns from a fixed dataset—a common requirement in realistic tasks to prevent unsafe exploration. To address this, we propose Diffusion-Regularized Constr…

Cited by 1SourcecodeScholar
2025

Subsampled Ensemble Can Improve Generalization Tail Exponentially

NeurIPS 2025poster

Ensemble learning is a popular technique to improve the accuracy of machine learning models. It traditionally hinges on the rationale that aggregating multiple weak models can lead to better models with lower variance and hence higher stability, especially for discontinuous base learners. In this pa…

Cited by 0SourcecodeScholar
2023

No-Regret Learning in Dynamic Competition with Reference Effects Under Logit Demand

NeurIPS 2023poster

This work is dedicated to the algorithm design in a competitive framework, with the primary goal of learning a stable equilibrium. We consider the dynamic price competition between two firms operating within an opaque marketplace, where each firm lacks information about its competitor. The demand fo…

Cited by 2SourcePDFScholar
2023

Policy-Based Primal-Dual Methods for Convex Constrained Markov Decision Processes

AAAI 2023technical

We study convex Constrained Markov Decision Processes (CMDPs) in which the objective is concave and the constraints are convex in the state-action occupancy measure. We propose a policy-based primal-dual algorithm that updates the primal variable via policy gradient ascent and updates the dual varia…

Cited by 15SourcePDFScholar
2023

Scalable Primal-Dual Actor-Critic Method for Safe Multi-Agent RL with General Utilities

NeurIPS 2023poster

We investigate safe multi-agent reinforcement learning, where agents seek to collectively maximize an aggregate sum of local objectives while satisfying their own safety constraints. The objective and constraints are described by general utilities, i.e., nonlinear functions of the long-term state-ac…

Cited by 16SourcePDFScholar
2022

A Dual Approach to Constrained Markov Decision Processes with Entropy Regularization

AISTATS 2022poster

We study entropy-regularized constrained Markov decision processes (CMDPs) under the soft-max parameterization, in which an agent aims to maximize the entropy-regularized value function while satisfying constraints on the expected total utility. By leveraging the entropy regularization, our theoreti…

Cited by 44SourcePDFScholar