← Search

Sinho Chewi

12 accepted papers

2026

High-accuracy sampling for diffusion models and log-concave distributions

ICML 2026oral

We present algorithms for diffusion model sampling which obtain $\delta$-error in $\mathrm{polylog}(1/\delta)$ steps, given access to $\widetilde O(\delta)$-accurate score estimates in $L^2$. This is an exponential improvement over all previous results. Specifically, under minimal data assumptions, …

Cited by 0SourceScholar
2023

Forward-Backward Gaussian Variational Inference via JKO in the Bures-Wasserstein Space

ICML 2023poster

Variational inference (VI) seeks to approximate a target distribution $\pi$ by an element of a tractable family of distributions. Of key interest in statistics and machine learning is Gaussian VI, which approximates $\pi$ by minimizing the Kullback-Leibler (KL) divergence to $\pi$ over the space of…

Cited by 43SourcePDFScholar
2023

Learning threshold neurons via edge of stability

NeurIPS 2023poster

Existing analyses of neural network training often operate under the unrealistic assumption of an extremely small learning rate. This lies in stark contrast to practical wisdom and empirical studies, such as the work of J. Cohen et al. (ICLR 2021), which exhibit startling new phenomena (the "edge of…

Cited by 47SourcePDFScholar
2023

Sampling is as easy as learning the score: theory for diffusion models with minimal data assumptions

ICLR 2023top-5%

We provide theoretical convergence guarantees for score-based generative models (SGMs) such as denoising diffusion probabilistic models (DDPMs), which constitute the backbone of large-scale real-world generative models such as DALL$\cdot$E 2. Our main result is that, assuming accurate score estimate…

Cited by 335SourcePDFScholar
2023

The probability flow ODE is provably fast

NeurIPS 2023poster

We provide the first polynomial-time convergence guarantees for the probabilistic flow ODE implementation (together with a corrector step) of score-based generative modeling. Our analysis is carried out in the wake of recent results obtaining such guarantees for the SDE-based implementation (i.e., d…

Cited by 166SourcePDFScholar
2022

Rejection sampling from shape-constrained distributions in sublinear time

AISTATS 2022poster

We consider the task of generating exact samples from a target distribution, known up to normalization, over a finite alphabet. The classical algorithm for this task is rejection sampling, and although it has been used in practice for decades, there is surprisingly little study of its fundamental li…

2022

Variational inference via Wasserstein gradient flows

NeurIPS 2022accept

Along with Markov chain Monte Carlo (MCMC) methods, variational inference (VI) has emerged as a central computational approach to large-scale Bayesian inference. Rather than sampling from the true posterior $\pi$, VI aims at producing a simple but effective approximation $\hat \pi$ to $\pi$ for whic…

2021

Averaging on the Bures-Wasserstein manifold: dimension-free convergence of gradient descent

NeurIPS 2021spotlight

We study first-order optimization algorithms for computing the barycenter of Gaussian distributions with respect to the optimal transport metric. Although the objective is geodesically non-convex, Riemannian gradient descent empirically converges rapidly, in fact faster than off-the-shelf methods su…

Cited by 60SourcePDFScholar
2021

Fast and Smooth Interpolation on Wasserstein Space

AISTATS 2021poster

We propose a new method for smoothly interpolating probability measures using the geometry of optimal transport. To that end, we reduce this problem to the classical Euclidean setting, allowing us to directly leverage the extensive toolbox of spline interpolation. Unlike previous approaches to measu…

Cited by 40SourcePDFScholar
2020

Exponential ergodicity of mirror-Langevin diffusions

NeurIPS 2020poster

Motivated by the problem of sampling from ill-conditioned log-concave distributions, we give a clean non-asymptotic convergence analysis of mirror-Langevin diffusions as introduced in Zhang et al. (2020). As a special case of this framework, we propose a class of diffusions called Newton-Langevin di…

Cited by 60SourcePDFScholar
2020

SVGD as a kernelized Wasserstein gradient flow of the chi-squared divergence

NeurIPS 2020poster

Stein Variational Gradient Descent (SVGD), a popular sampling algorithm, is often described as the kernelized gradient flow for the Kullback-Leibler divergence in the geometry of optimal transport. We introduce a new perspective on SVGD that instead views SVGD as the kernelized gradient flow of the…

Cited by 90SourcePDFScholar