← Search

Yu-Jie Zhang

16 accepted papers

2026

Dynamic Regret via Discounted-to-Dynamic Reduction with Applications to Curved Losses and Adam Optimizer

ICML 2026poster

We study dynamic regret minimization in non-stationary online learning, with a primary focus on follow-the-regularized-leader (FTRL) methods. FTRL is important for curved losses and for understanding adaptive algorithms, yet existing dynamic regret analyses are less explored for FTRL. To address thi…

Cited by 0SourceScholar
2025

Generalized Linear Bandits: Almost Optimal Regret with One-Pass Update

NeurIPS 2025poster

We study the generalized linear bandit (GLB) problem, a contextual multi-armed bandit framework that extends the classical linear model by incorporating a non-linear link function, thereby modeling a broad class of reward distributions such as Bernoulli and Poisson. While GLBs are widely applicable…

Cited by 0SourceScholar
2025

Heavy-Tailed Linear Bandits: Huber Regression with One-Pass Update

ICML 2025poster

We study the stochastic linear bandits with heavy-tailed noise. Two principled strategies for handling heavy-tailed noise, truncation and median-of-means, have been introduced to heavy-tailed bandits. Nonetheless, these methods rely on specific noise assumptions or bandit structures, limiting their…

Cited by 0SourcePDFScholar
2025

Non-stationary Online Learning for Curved Losses: Improved Dynamic Regret via Mixability

ICML 2025poster

Non-stationary online learning has drawn much attention in recent years. Despite considerable progress, dynamic regret minimization has primarily focused on convex functions, leaving the functions with stronger curvature (e.g., squared or logistic loss) underexplored. In this work, we address this g…

Cited by 0SourcePDFScholar
2024

Efficient Non-stationary Online Learning by Wavelets with Applications to Online Distribution Shift Adaptation

ICML 2024poster

Dynamic regret minimization offers a principled way for non-stationary online learning, where the algorithm's performance is evaluated against changing comparators. Prevailing methods often employ a two-layer online ensemble, consisting of a group of base learners with different configurations and a…

Cited by 2SourcePDFScholar
2024

Learning with Complementary Labels Revisited: The Selected-Completely-at-Random Setting Is More Practical

ICML 2024poster

Complementary-label learning is a weakly supervised learning problem in which each training example is associated with one or multiple complementary labels indicating the classes to which it does not belong. Existing consistent approaches have relied on the uniform distribution assumption to model t…

2024

Provably Efficient Reinforcement Learning with Multinomial Logit Function Approximation

NeurIPS 2024poster

We study a new class of MDPs that employs multinomial logit (MNL) function approximation to ensure valid probability distributions over the state space. Despite its significant benefits, incorporating the non-linear function raises substantial challenges in both *statistical* and *computational* eff…

Cited by 2SourcePDFScholar
2023

Adapting to Continuous Covariate Shift via Online Density Ratio Estimation

NeurIPS 2023poster

Dealing with distribution shifts is one of the central challenges for modern machine learning. One fundamental situation is the covariate shift, where the input distributions of data change from the training to testing stages while the input-conditional output distribution remains unchanged. In this…

Cited by 17SourcePDFScholar
2023

Online (Multinomial) Logistic Bandit: Improved Regret and Constant Computation Cost

NeurIPS 2023spotlight

This paper investigates the logistic bandit problem, a variant of the generalized linear bandit model that utilizes a logistic model to depict the feedback from an action. While most existing research focuses on the binary logistic bandit problem, the multinomial case, which considers more than two…

Cited by 18SourcePDFScholar
2022

Adapting to Online Label Shift with Provable Guarantees

NeurIPS 2022accept

The standard supervised learning paradigm works effectively when training data shares the same distribution as the upcoming testing samples. However, this stationary assumption is often violated in real-world applications, especially when testing data appear in an online fashion. In this paper, we f…

Cited by 36SourcePDFScholar
2020

An Unbiased Risk Estimator for Learning with Augmented Classes

NeurIPS 2020poster

This paper studies the problem of learning with augmented classes (LAC), where augmented classes unobserved in the training data might emerge in the testing phase. Previous studies generally attempt to discover augmented classes by exploiting geometric properties, achieving inspiring empirical perfo…

Cited by 31SourcePDFScholar