← Search

Jerry Li

35 accepted papers

2025

BEVCalib: LiDAR-Camera Calibration via Geometry-Guided Bird’s-Eye View Representation

CoRL 2025poster

Accurate LiDAR-camera calibration is the foundation of accurate multimodal fusion environmental perception for autonomous driving and robotic systems. Traditional calibration methods require extensive data collection in controlled environments and cannot compensate for the transformation changes dur…

Cited by 0SourceScholar
2025

S4S: Solving for a Fast Diffusion Model Solver

ICML 2025poster

Diffusion models (DMs) create samples from a data distribution by starting from random noise and iteratively solving a reverse-time ordinary differential equation (ODE). Because each step in the iterative solution requires an expensive neural function evaluation (NFE), there has been significant int…

Cited by 0SourcePDFScholar
2024

KITAB: Evaluating LLMs on Constraint Satisfaction for Information Retrieval

ICLR 2024poster

We study the ability of state-of-the art models to answer constraint satisfaction queries for information retrieval (e.g., “a list of ice cream shops in San Diego”). In the past, such queries were considered as tasks that could only be solved via web-search or knowledge bases. More recently, large l…

Cited by 10SourcePDFScholar
2024

Semi-Random Matrix Completion via Flow-Based Adaptive Reweighting

NeurIPS 2024poster

We consider the well-studied problem of completing a rank-$r$, $\mu$-incoherent matrix $\mathbf{M} \in \mathbb{R}^{d \times d}$ from incomplete observations. We focus on this problem in the semi-random setting where each entry is independently revealed with probability at least $p = \frac{\textup{po…

Cited by 0SourcePDFScholar
2023

Automatic Prompt Optimization with "Gradient Descent" and Beam Search

EMNLP 2023long main

Large Language Models (LLMs) have shown impressive performance as general purpose agents, but their abilities remain highly dependent on prompts which are hand written with onerous trial-and-error effort. We propose a simple and nonparametric solution to this problem, Prompt Optimization with Textua…

Cited by 0SourcecodeScholar
2023

Sampling is as easy as learning the score: theory for diffusion models with minimal data assumptions

ICLR 2023top-5%

We provide theoretical convergence guarantees for score-based generative models (SGMs) such as denoising diffusion probabilistic models (DDPMs), which constitute the backbone of large-scale real-world generative models such as DALL$\cdot$E 2. Our main result is that, assuming accurate score estimate…

Cited by 335SourcePDFScholar
2023

Structured Semidefinite Programming for Recovering Structured Preconditioners

NeurIPS 2023poster

We develop a general framework for finding approximately-optimal preconditioners for solving linear systems. Leveraging this framework we obtain improved runtimes for fundamental preconditioning and linear system solving problems including: Diagonal preconditioning. We give an algorithm which, given…

Cited by 5SourcePDFScholar
2022

Minimax Optimality (Probably) Doesn't Imply Distribution Learning for GANs

ICLR 2022poster

Arguably the most fundamental question in the theory of generative adversarial networks (GANs) is to understand when GANs can actually learn the underlying distribution. Theoretical and empirical evidence (see e.g. Arora-Risteski-Zhang '18) suggest local optimality of the empirical training objectiv…

Cited by 8SourcePDFScholar
2021

Aligning AI With Shared Human Values

ICLR 2021poster

We show how to assess a language model's knowledge of basic concepts of morality. We introduce the ETHICS dataset, a new benchmark that spans concepts in justice, well-being, duties, virtues, and commonsense morality. Models predict widespread moral judgments about diverse text scenarios. This requi…

2021

Byzantine-Resilient Non-Convex Stochastic Gradient Descent

ICLR 2021poster

We study adversary-resilient stochastic distributed optimization, in which $m$ machines can independently compute stochastic gradients, and cooperate to jointly optimize over their local objective functions. However, an $\alpha$-fraction of the machines are Byzantine, in that they may behave in arbi…

Cited by 89SourcePDFScholar
2021

List-Decodable Mean Estimation in Nearly-PCA Time

NeurIPS 2021spotlight

Robust statistics has traditionally focused on designing estimators tolerant to a minority of contaminated data. {\em List-decodable learning}~\cite{CharikarSV17} studies the more challenging regime where only a minority $\tfrac 1 k$ fraction of the dataset, $k \geq 2$, is drawn from the distributio…

Cited by 24SourcePDFScholar
2021

Robust Regression Revisited: Acceleration and Improved Estimation Rates

NeurIPS 2021poster

We study fast algorithms for statistical regression problems under the strong contamination model, where the goal is to approximately optimize a generalized linear model (GLM) given adversarially corrupted samples. Prior works in this line of research were based on the \emph{robust gradient descent}…

Cited by 23SourcePDFScholar
2020

Learning Structured Distributions From Untrusted Batches: Faster and Simpler

NeurIPS 2020poster

We revisit the problem of learning from untrusted batches introduced by Qiao and Valiant [QV17]. Recently, Jain and Orlitsky [JO19] gave a simple semidefinite programming approach based on the cut-norm that achieves essentially information-theoretically optimal error in polynomial time. Concurrently…

2020

Low-Rank Toeplitz Matrix Estimation Via Random Ultra-Sparse Rulers

ICASSP 2020accepted

We study how to estimate a nearly low-rank Toeplitz covariance matrix T from compressed measurements. Recent work of Qiao and Pal addresses this problem by combining sparse rulers (sparse linear arrays) with frequency finding (sparse Fourier transform) algorithms applied to the Vandermonde decomposi…

Cited by 0SourceScholar
2020

RL Unplugged: A Suite of Benchmarks for Offline Reinforcement Learning

NeurIPS 2020poster

Offline methods for reinforcement learning have a potential to help bridge the gap between reinforcement learning research and real-world applications. They make it possible to learn policies from offline datasets, thus overcoming concerns associated with online data collection in the real-world, in…

2020

Randomized Smoothing of All Shapes and Sizes

ICML 2020poster

Randomized smoothing is the current state-of-the-art defense with provable robustness against $\ell_2$ adversarial attacks. Many works have devised new randomized smoothing schemes for other metrics, such as $\ell_1$ or $\ell_\infty$; however, substantial effort was needed to derive such new guarant…

2020

Robust Sub-Gaussian Principal Component Analysis and Width-Independent Schatten Packing

NeurIPS 2020spotlight

We develop two methods for the following fundamental statistical task: given an $\eps$-corrupted set of $n$ samples from a $d$-dimensional sub-Gaussian distribution, return an approximate top eigenvector of the covariance matrix. Our first robust PCA algorithm runs in polynomial time, returns a $1 -…

Cited by 53SourcePDFScholar
2019

Provably Robust Deep Learning via Adversarially Trained Smoothed Classifiers

NeurIPS 2019spotlight

Recent works have shown the effectiveness of randomized smoothing as a scalable technique for building neural network-based classifiers that are provably robust to $\ell_2$-norm adversarial perturbations. In this paper, we employ adversarial training to improve the performance of randomized smoothin…

2019

Quantum Entropy Scoring for Fast Robust Mean Estimation and Improved Outlier Detection

NeurIPS 2019spotlight

We study two problems in high-dimensional robust statistics: \emph{robust mean estimation} and \emph{outlier detection}. In robust mean estimation the goal is to estimate the mean $\mu$ of a distribution on $\mathbb{R}^d$ given $n$ independent samples, an $\epsilon$-fraction of which have been corru…

2019

Sever: A Robust Meta-Algorithm for Stochastic Optimization

ICML 2019oral

In high dimensions, most machine learning methods are brittle to even a small fraction of structured outliers. To address this, we introduce a new meta-algorithm that can take in a base learner such as least squares or stochastic gradient descent, and harden the learner to be resistant to outliers.…

2018

On the Limitations of First-Order Approximation in GAN Dynamics

ICML 2018oral

While Generative Adversarial Networks (GANs) have demonstrated promising performance on multiple vision tasks, their learning dynamics are not yet well understood, both in theory and in practice. To address this issue, we study GAN dynamics in a simple yet rich parametric model that exhibits several…

Cited by 69SourcePDFScholar
2018

On the limitations of first order approximation in GAN dynamics

ICLR 2018workshop

Generative Adversarial Networks (GANs) have been proposed as an approach to learning generative models. While GANs have demonstrated promising performance on multiple vision tasks, their learning dynamics are not yet well understood, neither in theory nor in practice. In particular, the work in this…

Cited by 69SourceScholar
2017

Being Robust (in High Dimensions) Can Be Practical

ICML 2017poster

Robust estimation is much more challenging in high-dimensions than it is in one-dimension: Most techniques either lead to intractable optimization problems or estimators that can tolerate only a tiny fraction of errors. Recent work in theoretical computer science has shown that, in appropriate distr…

2017

Communication-Efficient Distributed Learning of Discrete Distributions

NeurIPS 2017oral

We initiate a systematic investigation of distribution learning (density estimation) when the data is distributed across multiple servers. The servers must communicate with a referee and the goal is to estimate the underlying distribution with as few bits of communication as possible. We focus on no…

Cited by 49SourcePDFScholar
2017

QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding

NeurIPS 2017poster

Parallel implementations of stochastic gradient descent (SGD) have received significant research attention, thanks to its excellent scalability properties. A fundamental barrier when parallelizing SGD is the high bandwidth cost of communicating gradient updates between nodes; consequently, several l…

2017

ZipML: Training Linear Models with End-to-End Low Precision, and a Little Bit of Deep Learning

ICML 2017poster

Recently there has been significant interest in training machine-learning models at low precision: by reducing precision, one can reduce computation and communication by one order of magnitude. We examine training at reduced precision, both from a theoretical and practical perspective, and ask: is i…

Cited by 227SourcePDFScholar