← Search

Yulian Wu

7 accepted papers

2025

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO

ICML 2025spotlight

In this paper, we theoretically investigate the effects of noisy labels in offline alignment, with a focus on the interplay between privacy and robustness against adversarial corruption. Specifically, under linear modeling assumptions, we present a unified analysis covering both reinforcement learni…

Cited by 0SourcePDFScholar
2025

Optimal Regret of Bandits under Differential Privacy

NeurIPS 2025poster

As sequential learning algorithms are increasingly applied to real life, ensuring data privacy while maintaining their utilities emerges as a timely question. In this context, regret minimisation in stochastic bandits under $\epsilon$-global Differential Privacy (DP) has been widely studied. The pr…

Cited by 0SourceScholar
2025

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment

ICML 2025poster

In this paper, we theoretically study the offline alignment of language models with human preference feedback, under both preference label corruption and privacy protections. To this end, we propose a variant of \texttt{$\chi$PO} -- \texttt{Square}\texttt{$\chi$PO}, which is a simple one-line change…

Cited by 0SourcePDFScholar
2023

Differentially Private Episodic Reinforcement Learning with Heavy-tailed Rewards

ICML 2023poster

In this paper we study the problem of (finite horizon tabular) Markov decision processes (MDPs) with heavy-tailed rewards under the constraint of differential privacy (DP). Compared with the previous studies for private reinforcement learning that typically assume rewards are sampled from some bound…

Cited by 1SourcePDFScholar
2022

Optimal Rates of (Locally) Differentially Private Heavy-tailed Multi-Armed Bandits

AISTATS 2022poster

In this paper we investigate the problem of stochastic multi-armed bandits (MAB) in the (local) differential privacy (DP/LDP) model. Unlike previous results that assume bounded/sub-Gaussian reward distributions, we focus on the setting where each arm’s reward distribution only has $(1+v)$-th moment…

Cited by 38SourcePDFScholar
2022

Private Stochastic Convex Optimization and Sparse Learning with Heavy-tailed Data Revisited

IJCAI 2022poster

In this paper, we revisit the problem of Differentially Private Stochastic Convex Optimization (DP-SCO) with heavy-tailed data, where the gradient of the loss function has bounded moments. Instead of the case where the loss function is Lipschitz or each coordinate of the gradient has bounded second…

Cited by 13SourcePDFScholar