← Search

Zhiming Huang

2 accepted papers

2025

Connecting Thompson Sampling and UCB: Towards More Efficient Trade-offs Between Privacy and Regret

ICML 2025poster

We address differentially private stochastic bandit problems by leveraging Thompson Sampling with Gaussian priors and Gaussian differential privacy (GDP). We propose DP-TS-UCB, a novel parametrized private algorithm that enables trading off privacy and regret. DP-TS-UCB satisfies $ \tilde{O} \l…

Cited by 0SourcePDFScholar
2023

A near-optimal high-probability swap-Regret upper bound for multi-agent bandits in unknown general-sum games

UAI 2023poster

In this paper, we study a multi-agent bandit problem in an unknown general-sum game repeated for a number of rounds (i.e., learning in a black-box game with bandit feedback), where a set of agents have no information about the underlying game structure and cannot observe each other’s actions and rew…

Cited by 4SourcePDFScholar