← Search

Dongsheng Ding

11 accepted papers

2026

Unlearning in Diffusion Models: A Unified Framework with KL Divergence and Likelihood Constraints

ICML 2026poster

Unlearning in diffusion models aims to remove undesirable data or concepts while preserving the utility of pretrained models---two fundamentally conflicting objectives. We propose a principled constrained optimization framework that formulates unlearning as minimizing the deviation from a pretrained…

Cited by 0SourceScholar
2025

Alignment of Large Language Models with Constrained Learning

NeurIPS 2025poster

We study the problem of computing an optimal large language model (LLM) policy for the constrained alignment problem, where the goal is to maximize a primary reward objective while satisfying constraints on secondary utilities. Despite the popularity of Lagrangian-based LLM policy search in constrai…

Cited by 0SourceScholar
2025

Composition and Alignment of Diffusion Models using Constrained Learning

NeurIPS 2025poster

Diffusion models have become prevalent in generative modeling due to their ability to sample from complex distributions. To improve the quality of generated samples and their compliance with user requirements, two commonly used methods are: (i) Alignment, which involves finetuning a diffusion model…

Cited by 0SourcecodeScholar
2025

Deterministic Policy Gradient Primal-Dual Methods for Continuous-Space Constrained MDPs

AAAI 2025technical

We study the problem of computing deterministic optimal policies for constrained Markov decision processes (MDPs) with continuous state and action spaces, which are widely encountered in constrained dynamical systems. Designing deterministic policy gradient methods in continuous state and action spa…

2024

One-Shot Safety Alignment for Large Language Models via Optimal Dualization

NeurIPS 2024spotlight

The growing safety concerns surrounding large language models raise an urgent need to align them with diverse human preferences to simultaneously enhance their helpfulness and safety. A promising approach is to enforce safety constraints through Reinforcement Learning from Human Feedback (RLHF). For…

2023

Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs

NeurIPS 2023poster

We study the problem of computing an optimal policy of an infinite-horizon discounted constrained Markov decision process (constrained MDP). Despite the popularity of Lagrangian-based policy search methods used in practice, the oscillation of policy iterates in these methods has not been fully under…

Cited by 30SourcePDFScholar
2022

Independent Policy Gradient for Large-Scale Markov Potential Games: Sharper Rates, Function Approximation, and Game-Agnostic Convergence

ICML 2022oral

We examine global non-asymptotic convergence properties of policy gradient methods for multi-agent reinforcement learning (RL) problems in Markov potential games (MPGs). To learn a Nash equilibrium of an MPG in which the size of state space and/or the number of players can be very large, we propose…

Cited by 97SourcePDFScholar
2021

Provably Efficient Safe Exploration via Primal-Dual Policy Optimization

AISTATS 2021poster

We study the safe reinforcement learning problem using the constrained Markov decision processes in which an agent aims to maximize the expected total reward subject to a safety constraint on the expected total value of a utility function. We focus on an episodic setting with the function approximat…

Cited by 200SourcePDFScholar
2020

Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision Processes

NeurIPS 2020poster

We study sequential decision-making problems in which each agent aims to maximize the expected total reward while satisfying a constraint on the expected total utility. We employ the natural policy gradient method to solve the discounted infinite-horizon Constrained Markov Decision Processes (CMDPs)…

Cited by 241SourcePDFScholar