← Search

Wei Biao Wu

6 accepted papers

2026

Sharp asymptotic theory for Q-learning with \texttt{LD2Z} learning rate and its generalization

ICLR 2026poster

Despite the sustained popularity of Q-learning as a practical tool for policy determination, a majority of relevant theoretical literature deals with either constant ($\eta_t\equiv \eta$) or polynomially decaying ($\eta_t = \eta t^{-\alpha}$) learning schedules. However, it is well known the these c…

Cited by 0SourceScholar
2026

Stability beyond bounded differences: sharp generalization bounds under finite $L_p$ moments

ICML 2026poster

While algorithmic stability is a central tool for understanding generalization of learning algorithms, existing high-probability guarantees typically rely on uniform boundedness or sub-Gaussian/sub-Weibull tail assumptions, which can be overly restrictive for modern settings with heavy-tailed or unb…

Cited by 0SourceScholar
2025

Asymptotic theory of SGD with a general learning-rate

NeurIPS 2025poster

Stochastic gradient descent (SGD) with polynomially decaying step‐sizes has long underpinned theoretical analyses, yielding a broad spectrum of statistically attractive guarantees. Yet in practice, such schedules find rare use due to their prohibitively slow convergence, revealing a persistent gap b…

Cited by 0SourceScholar
2025

Gaussian Approximation and Concentration of Constant Learning-Rate Stochastic Gradient Descent

NeurIPS 2025poster

We establish a comprehensive finite-sample and asymptotic theory for stochastic gradient descent (SGD) with constant learning rates. First, we propose a novel linear approximation technique to provide a quenched central limit theorem (CLT) for SGD iterates with refined tail properties, showing that…

Cited by 0SourceScholar
2025

Statistical Guarantees for High-Dimensional Stochastic Gradient Descent

NeurIPS 2025poster

Stochastic Gradient Descent (SGD) and its Ruppert–Polyak averaged variant (ASGD) lie at the heart of modern large-scale learning, yet their theoretical properties in high-dimensional settings are rarely understood. In this paper, we provide rigorous statistical guarantees for constant learning-rate…

Cited by 0SourceScholar