← Search

Yi Hao

12 accepted papers

2022

TURF: Two-Factor, Universal, Robust, Fast Distribution Learning Algorithm

ICML 2022spotlight

Approximating distributions from their samples is a canonical statistical-learning problem. One of its most powerful and successful modalities approximates every distribution to an $\ell_1$ distance essentially at most a constant times larger than its closest $t$-piece degree-$d$ polynomial, where $…

Cited by 0SourcePDFScholar
2021

Compressed Maximum Likelihood

ICML 2021spotlight

Maximum likelihood (ML) is one of the most fundamental and general statistical estimation techniques. Inspired by recent advances in estimating distribution functionals, we propose $\textit{compressed maximum likelihood}$ (CML) that applies ML to the compressed samples. We then show that CML is samp…

Cited by 0SourcePDFScholar
2020

Profile Entropy: A Fundamental Measure for the Learnability and Compressibility of Distributions

NeurIPS 2020poster

The profile of a sample is the multiset of its symbol frequencies. We show that for samples of discrete distributions, profile entropy is a fundamental measure unifying the concepts of estimation, inference, and compression. Specifically, profile entropy: a) determines the speed of estimating the d…

Cited by 7SourcePDFScholar
2020

SURF: A Simple, Universal, Robust, Fast Distribution Learning Algorithm

NeurIPS 2020poster

Sample- and computationally-efficient distribution estimation is a fundamental tenet in statistics and machine learning. We present $\SURF$, an algorithm for approximating distributions by piecewise polynomials. $\SURF$ is: simple, replacing prior complex optimization techniques by straight-forward…

Cited by 9SourcePDFScholar
2018

Data Amplification: A Unified and Competitive Approach to Property Estimation

NeurIPS 2018poster

Estimating properties of discrete distributions is a fundamental problem in statistical learning. We design the first unified, linear-time, competitive, property estimator that for a wide class of properties and for all underlying distributions uses just 2n samples to achieve the performance attaine…

Cited by 30SourcePDFScholar