← Search

Jasper C.H. Lee

14 accepted papers

2025

All-Purpose Mean Estimation over R: Optimal Sub-Gaussianity with Outlier Robustness and Low Moments Performance

ICML 2025oral

We consider the basic statistical challenge of designing an "all-purpose" mean estimation algorithm that is recommendable across a variety of settings and models. Recent work by [Lee and Valiant 2022] introduced the first 1-d mean estimator whose error in the standard finite-variance+i.i.d. setting…

Cited by 0SourcePDFScholar
2025

On Fine-Grained Distinct Element Estimation

ICML 2025poster

We study the problem of distributed distinct element estimation, where $\alpha$ servers each receive a subset of a universe $[n]$ and aim to compute a $(1+\varepsilon)$-approximation to the number of distinct elements using minimal communication. While prior work establishes a worst-case bound of $\…

Cited by 0SourcePDFScholar
2025

On Learning Parallel Pancakes with Mostly Uniform Weights

ICML 2025spotlight

We study the complexity of learning $k$-mixtures of Gaussians ($k$-GMMs) on $\mathbb R^d$. This task is known to have complexity $d^{\Omega(k)}$ in full generality. To circumvent this exponential lower bound on the number of components, research has focused on learning families of GMMs satisfying ad…

Cited by 0SourcePDFScholar
2024

Multi-Stage Predict+Optimize for (Mixed Integer) Linear Programs

NeurIPS 2024poster

The recently-proposed framework of Predict+Optimize tackles optimization problems with parameters that are unknown at solving time, in a supervised learning setting. Prior frameworks consider only the scenario where all unknown parameters are (eventually) revealed simultaneously. In this work, we pr…

Cited by 0SourcePDFScholar
2023

A Spectral Algorithm for List-Decodable Covariance Estimation in Relative Frobenius Norm

NeurIPS 2023spotlight

We study the problem of list-decodable Gaussian covariance estimation. Given a multiset $T$ of $n$ points in $\mathbb{R}^d$ such that an unknown $\alpha<1/2$ fraction of points in $T$ are i.i.d. samples from an unknown Gaussian $\mathcal{N}(\mu, \Sigma)$, the goal is to output a list of $O(1/\alpha)…

Cited by 4SourcePDFScholar
2023

High-dimensional Location Estimation via Norm Concentration for Subgamma Vectors

ICML 2023poster

In location estimation, we are given $n$ samples from a known distribution $f$ shifted by an unknown translation $\lambda$, and want to estimate $\lambda$ as precisely as possible. Asymptotically, the maximum likelihood estimate achieves the Cramér-Rao bound of error $\mathcal N(0, \frac{1}{n\mathca…

Cited by 7SourcePDFScholar
2023

Optimality in Mean Estimation: Beyond Worst-Case, Beyond Sub-Gaussian, and Beyond $1+\alpha$ Moments

NeurIPS 2023poster

There is growing interest in improving our algorithmic understanding of fundamental statistical problems such as mean estimation, driven by the goal of understanding the fundamental limits of what we can extract from limited and valuable data. The state of the art results for mean estimation in $\ma…

Cited by 2SourcePDFScholar
2023

Predict+Optimize for Packing and Covering LPs with Unknown Parameters in Constraints

AAAI 2023technical

Predict+Optimize is a recently proposed framework which combines machine learning and constrained optimization, tackling optimization problems that contain parameters that are unknown at solving time. The goal is to predict the unknown parameters and use the estimates to solve for an estimated optim…

Cited by 16SourcePDFScholar
2023

Two-Stage Predict+Optimize for MILPs with Unknown Parameters in Constraints

NeurIPS 2023poster

Consider the setting of constrained optimization, with some parameters unknown at solving time and requiring prediction from relevant features. Predict+Optimize is a recent framework for end-to-end training supervised learning models for such predictions, incorporating information about the optimiza…

Cited by 10SourcePDFScholar
2022

Branch & Learn for Recursively and Iteratively Solvable Problems in Predict+Optimize

NeurIPS 2022accept

This paper proposes Branch & Learn, a framework for Predict+Optimize to tackle optimization problems containing parameters that are unknown at the time of solving. Given an optimization problem solvable by a recursive algorithm satisfying simple conditions, we show how a corresponding learning algor…

Cited by 10SourcePDFScholar
2022

Finite-Sample Maximum Likelihood Estimation of Location

NeurIPS 2022accept

We consider 1-dimensional location estimation, where we estimate a parameter $\lambda$ from $n$ samples $\lambda + \eta_i$, with each $\eta_i$ drawn i.i.d. from a known distribution $f$. For fixed $f$ the maximum-likelihood estimate (MLE) is well-known to be optimal in the limit as $n \to \infty$: i…

Cited by 9SourcePDFScholar
2022

Outlier-Robust Sparse Mean Estimation for Heavy-Tailed Distributions

NeurIPS 2022accept

We study the fundamental task of outlier-robust mean estimation for heavy-tailed distributions in the presence of sparsity. Specifically, given a small number of corrupted samples from a high-dimensional heavy-tailed distribution whose mean $\mu$ is guaranteed to be sparse, the goal is to efficient…

Cited by 16SourcePDFScholar
2021

Quantifying and Reducing Bias in Maximum Likelihood Estimation of Structured Anomalies

ICML 2021spotlight

Anomaly estimation, or the problem of finding a subset of a dataset that differs from the rest of the dataset, is a classic problem in machine learning and data mining. In both theoretical work and in applications, the anomaly is assumed to have a specific structure defined by membership in an anoma…

Cited by 7SourcePDFScholar