← Search

Yan Dai

10 accepted papers

2025

Incentive-Aware Dynamic Resource Allocation under Long-Term Cost Constraints

NeurIPS 2025poster

Motivated by applications such as cloud platforms allocating GPUs to users or governments deploying mobile health units across competing regions, we study the constrained dynamic allocation of a reusable resource to a group of strategic agents. Our objective is to simultaneously (i) maximize social…

Cited by 0SourceScholar
2025

uniINF: Best-of-Both-Worlds Algorithm for Parameter-Free Heavy-Tailed MABs

ICLR 2025spotlight

In this paper, we present a novel algorithm, `uniINF`, for the Heavy-Tailed Multi-Armed Bandits (HTMAB) problem, demonstrating robustness and adaptability in both stochastic and adversarial environments. Unlike the stochastic MAB setting where loss distributions are stationary with time, our study e…

Cited by 0SourcePDFScholar
2024

Understanding Adam Optimizer via Online Learning of Updates: Adam is FTRL in Disguise

ICML 2024poster

Despite the success of the Adam optimizer in practice, the theoretical understanding of its algorithmic components still remains limited. In particular, most existing analyses of Adam show the convergence rate that can be simply achieved by non-adative algorithms like SGD. In this work, we provide a…

Cited by 13SourcePDFScholar
2023

Banker Online Mirror Descent: A Universal Approach for Delayed Online Bandit Learning

ICML 2023poster

We propose Banker Online Mirror Descent (Banker-OMD), a novel framework generalizing the classical Online Mirror Descent (OMD) technique in the online learning literature. The Banker-OMD framework almost completely decouples feedback delay handling and the task-specific OMD algorithm design, thus fa…

Cited by 6SourcePDFScholar
2023

Refined Regret for Adversarial MDPs with Linear Function Approximation

ICML 2023poster

We consider learning in an adversarial Markov Decision Process (MDP) where the loss functions can change arbitrarily over $K$ episodes and the state space can be arbitrarily large. We assume that the Q-function of any policy is linear in some known features, that is, a linear function approximation…

Cited by 25SourcePDFScholar
2022

Adaptive Best-of-Both-Worlds Algorithm for Heavy-Tailed Multi-Armed Bandits

ICML 2022spotlight

In this paper, we generalize the concept of heavy-tailed multi-armed bandits to adversarial environments, and develop robust best-of-both-worlds algorithms for heavy-tailed multi-armed bandits (MAB), where losses have $\alpha$-th ($1<\alpha\le 2$) moments bounded by $\sigma^\alpha$, while the varian…

Cited by 20SourcePDFScholar
2022

Follow-the-Perturbed-Leader for Adversarial Markov Decision Processes with Bandit Feedback

NeurIPS 2022accept

We consider regret minimization for Adversarial Markov Decision Processes (AMDPs), where the loss functions are changing over time and adversarially chosen, and the learner only observes the losses for the visited state-action pairs (i.e., bandit feedback). While there has been a surge of studies on…

Cited by 18SourcePDFScholar
2021

RSGNet: Relation based Skeleton Graph Network for Crowded Scenes Pose Estimation

AAAI 2021technical

Despite of the recent great progress on multi-person pose estimation, existing solutions still remain challenging under the condition of "crowded scenes'', where RGB images capture complex real-world scenes with highly-overlapped people, severe occlusions and diverse postures. In this work, we focu…