← Search

Holden Lee

17 accepted papers

2024

How Flawed Is ECE? An Analysis via Logit Smoothing

ICML 2024poster

Informally, a model is calibrated if its predictions are correct with a probability that matches the confidence of the prediction. By far the most common method in the literature for measuring calibration is the expected calibration error (ECE). Recent work, however, has pointed out drawbacks of ECE…

2024

Principled Gradient-Based MCMC for Conditional Sampling of Text

ICML 2024poster

We consider the problem of sampling text from an energy-based model. This arises, for example, when sampling text from a neural language model subject to soft constraints. Although the target distribution is discrete, the internal computations of the energy function (given by the language model) are…

Cited by 1SourcePDFScholar
2024

What does guidance do? A fine-grained analysis in a simple setting

NeurIPS 2024poster

The use of guidance in diffusion models was originally motivated by the premise that the guidance-modified score is that of the data distribution tilted by a conditional likelihood raised to some power. In this work we clarify this misconception by rigorously proving that guidance fails to sample fr…

Cited by 10SourcePDFScholar
2023

Connecting Pre-trained Language Model and Downstream Task via Properties of Representation

NeurIPS 2023poster

Recently, researchers have found that representations learned by large-scale pre-trained language models are useful in various downstream tasks. However, there is little theoretical understanding of how pre-training performance is related to downstream task performance. In this paper, we analyze how…

Cited by 0SourcePDFScholar
2023

Improved Analysis of Score-based Generative Modeling: User-Friendly Bounds under Minimal Smoothness Assumptions

ICML 2023poster

We give an improved theoretical analysis of score-based generative modeling. Under a score estimate with small $L^2$ error (averaged across timesteps), we provide efficient convergence guarantees for any data distribution with second-order moment, by either employing early stopping or assuming smoot…

Cited by 182SourcePDFScholar
2023

Pitfalls of Gaussians as a noise distribution in NCE

ICLR 2023poster

Noise Contrastive Estimation (NCE) is a popular approach for learning probability density functions parameterized up to a constant of proportionality. The main idea is to design a classification problem for distinguishing training data from samples from an (easy-to-sample) noise distribution $q$, in…

Cited by 7SourcePDFScholar
2023

Provable benefits of score matching

NeurIPS 2023spotlight

Score matching is an alternative to maximum likelihood (ML) for estimating a probability distribution parametrized up to a constant of proportionality. By fitting the ''score'' of the distribution, it sidesteps the need to compute this constant of proportionality (which is often intractable). While…

Cited by 16SourcePDFScholar
2023

The probability flow ODE is provably fast

NeurIPS 2023poster

We provide the first polynomial-time convergence guarantees for the probabilistic flow ODE implementation (together with a corrector step) of score-based generative modeling. Our analysis is carried out in the wake of recent results obtaining such guarantees for the SDE-based implementation (i.e., d…

Cited by 166SourcePDFScholar
2022

Convergence for score-based generative modeling with polynomial complexity

NeurIPS 2022accept

Score-based generative modeling (SGM) is a highly successful approach for learning a probability distribution from data and generating further samples. We prove the first polynomial convergence guarantees for the core mechanic behind SGM: drawing samples from a probability density $p$ given a score…

Cited by 163SourcePDFScholar
2022

Extracting Latent State Representations with Linear Dynamics from Rich Observations

ICML 2022spotlight

Recently, many reinforcement learning techniques have been shown to have provable guarantees in the simple case of linear dynamics, especially in problems like linear quadratic regulators. However, in practice many tasks require learning a policy from rich, high-dimensional features such as images,…

Cited by 2SourcePDFScholar
2021

Universal Approximation Using Well-Conditioned Normalizing Flows

NeurIPS 2021poster

Normalizing flows are a widely used class of latent-variable generative models with a tractable likelihood. Affine-coupling models [Dinh et al., 2014, 2016] are a particularly common type of normalizing flows, for which the Jacobian of the latent-to-observable-variable transformation is triangular,…

Cited by 20SourcePDFScholar
2019

Explaining Landscape Connectivity of Low-cost Solutions for Multilayer Nets

NeurIPS 2019poster

Mode connectivity is a surprising phenomenon in the loss landscape of deep nets. Optima---at least those discovered by gradient-based optimization---turn out to be connected by simple paths on which the loss function is almost constant. Often, these paths can be chosen to be piece-wise linear, with…

2018

Beyond Log-concavity: Provable Guarantees for Sampling Multi-modal Distributions using Simulated Tempering Langevin Monte Carlo

NeurIPS 2018poster

A key task in Bayesian machine learning is sampling from distributions that are only specified up to a partition function (i.e., constant of proportionality). One prevalent example of this is sampling posteriors in parametric distributions, such as latent-variable generative models. However sampli…

Cited by 53SourcePDFScholar
2018

Spectral Filtering for General Linear Dynamical Systems

NeurIPS 2018oral

We give a polynomial-time algorithm for learning latent-state linear dynamical systems without system identification, and without assumptions on the spectral radius of the system's transition matrix. The algorithm extends the recently introduced technique of spectral filtering, previously applied on…

Cited by 115SourcePDFScholar
2018

Towards Provable Control for Unknown Linear Dynamical Systems

ICLR 2018workshop

We study the control of symmetric linear dynamical systems with unknown dynamics and a hidden state. Using a recent spectral filtering technique for concisely representing such systems in a linear basis, we formulate optimal control in this setting as a convex program. This approach eliminates the n…

Cited by 29SourceScholar