← Search

Lingjiong Zhu

11 accepted papers

2025

Convergence Analysis for General Probability Flow ODEs of Diffusion Models in Wasserstein Distances

AISTATS 2025poster

Score-based generative modeling with probability flow ordinary differential equations (ODEs) has achieved remarkable success in a variety of applications. While various fast ODE-based samplers have been proposed in the literature and employed in practice, the theoretical understandings about converg…

Cited by 0SourceScholar
2025

Go With the Flow: Fast Diffusion for Gaussian Mixture Models

NeurIPS 2025spotlight

Schrodinger Bridges (SBs) are diffusion processes that steer, in finite time, a given initial distribution to another final one while minimizing a suitable cost functional. Although various methods for computing SBs have recently been proposed in the literature, most of these approaches require com…

Cited by 0SourcecodeScholar
2023

Algorithmic Stability of Heavy-Tailed SGD with General Loss Functions

ICML 2023poster

Heavy-tail phenomena in stochastic gradient descent (SGD) have been reported in several empirical studies. Experimental evidence in previous works suggests a strong interplay between the heaviness of the tails and generalization behavior of SGD. To address this empirical phenomena theoretically, sev…

Cited by 26SourcePDFScholar
2023

Uniform-in-Time Wasserstein Stability Bounds for (Noisy) Stochastic Gradient Descent

NeurIPS 2023poster

Algorithmic stability is an important notion that has proven powerful for deriving generalization bounds for practical algorithms. The last decade has witnessed an increasing number of stability bounds for different algorithms applied on different classes of loss functions. While these bounds have i…

Cited by 9SourcePDFScholar
2021

Asymmetric Heavy Tails and Implicit Bias in Gaussian Noise Injections

ICML 2021spotlight

Gaussian noise injections (GNIs) are a family of simple and widely-used regularisation methods for training neural networks, where one injects additive or multiplicative Gaussian noise to the network activations at every iteration of the optimisation algorithm, which is typically chosen as stochasti…

2021

Convergence Rates of Stochastic Gradient Descent under Infinite Noise Variance

NeurIPS 2021poster

Recent studies have provided both empirical and theoretical evidence illustrating that heavy tails can emerge in stochastic gradient descent (SGD) in various scenarios. Such heavy tails potentially result in iterates with diverging variance, which hinders the use of conventional convergence analysis…

Cited by 51SourcePDFScholar
2021

Fractal Structure and Generalization Properties of Stochastic Optimization Algorithms

NeurIPS 2021spotlight

Understanding generalization in deep learning has been one of the major challenges in statistical learning theory over the last decade. While recent work has illustrated that the dataset and the training algorithm must be taken into account in order to obtain meaningful generalization bounds, it is…

Cited by 31SourcePDFScholar
2020

Breaking Reversibility Accelerates Langevin Dynamics for Non-Convex Optimization

NeurIPS 2020poster

Langevin dynamics (LD) has been proven to be a powerful technique for optimizing a non-convex objective as an efficient algorithm to find local minima while eventually visiting a global minimum on longer time-scales. LD is based on the first-order Langevin diffusion which is reversible in time. We s…

Cited by 21SourcePDFScholar
2020

Fractional Underdamped Langevin Dynamics: Retargeting SGD with Momentum under Heavy-Tailed Gradient Noise

ICML 2020poster

Stochastic gradient descent with momentum (SGDm) is one of the most popular optimization algorithms in deep learning. While there is a rich theory of SGDm for convex problems, the theory is considerably less developed in the context of deep learning where the problem is non-convex and the gradient n…

2019

Accelerated Linear Convergence of Stochastic Momentum Methods in Wasserstein Distances

ICML 2019oral

Momentum methods such as Polyak’s heavy ball (HB) method, Nesterov’s accelerated gradient (AG) as well as accelerated projected gradient (APG) method have been commonly used in machine learning practice, but their performance is quite sensitive to noise in the gradients. We study these methods under…

Cited by 54SourcePDFScholar