← Search

Philippe Rigollet

20 accepted papers

2023

The emergence of clusters in self-attention dynamics

NeurIPS 2023poster

Viewing Transformers as interacting particle systems, we describe the geometry of learned representations when the weights are not time-dependent. We show that particles, representing tokens, tend to cluster toward particular limiting objects as time tends to infinity. Using techniques from dynamica…

2022

GULP: a prediction-based metric between representations

NeurIPS 2022accept

Comparing the representations learned by different neural networks has recently emerged as a key tool to understand various architectures and ultimately optimize them. In this work, we introduce GULP, a family of distance measures between representations that is explicitly motivated by downstream p…

2022

Rejection sampling from shape-constrained distributions in sublinear time

AISTATS 2022poster

We consider the task of generating exact samples from a target distribution, known up to normalization, over a finite alphabet. The classical algorithm for this task is rejection sampling, and although it has been used in practice for decades, there is surprisingly little study of its fundamental li…

2022

Variational inference via Wasserstein gradient flows

NeurIPS 2022accept

Along with Markov chain Monte Carlo (MCMC) methods, variational inference (VI) has emerged as a central computational approach to large-scale Bayesian inference. Rather than sampling from the true posterior $\pi$, VI aims at producing a simple but effective approximation $\hat \pi$ to $\pi$ for whic…

2021

Fast and Smooth Interpolation on Wasserstein Space

AISTATS 2021poster

We propose a new method for smoothly interpolating probability measures using the geometry of optimal transport. To that end, we reduce this problem to the classical Euclidean setting, allowing us to directly leverage the extensive toolbox of spline interpolation. Unlike previous approaches to measu…

Cited by 40SourcePDFScholar
2020

Exponential ergodicity of mirror-Langevin diffusions

NeurIPS 2020poster

Motivated by the problem of sampling from ill-conditioned log-concave distributions, we give a clean non-asymptotic convergence analysis of mirror-Langevin diffusions as introduced in Zhang et al. (2020). As a special case of this framework, we propose a class of diffusions called Newton-Langevin di…

Cited by 60SourcePDFScholar
2020

SVGD as a kernelized Wasserstein gradient flow of the chi-squared divergence

NeurIPS 2020poster

Stein Variational Gradient Descent (SVGD), a popular sampling algorithm, is often described as the kernelized gradient flow for the Kullback-Leibler divergence in the geometry of optimal transport. We introduce a new perspective on SVGD that instead views SVGD as the kernelized gradient flow of the…

Cited by 90SourcePDFScholar
2019

Statistical Optimal Transport via Factored Couplings

AISTATS 2019poster

We propose a new method to estimate Wasserstein distances and optimal transport plans between two probability distributions from samples in high dimension. Unlike plug-in rules that simply replace the true distributions by their empirical counterparts, our method promotes couplings with low transpor…

Cited by 89SourcePDFScholar
2018

Teacher Improves Learning by Selecting a Training Subset

AISTATS 2018poster

We call a learner super-teachable if a teacher can trim down an iid training set while making the learner learn even better. We provide sharp super-teaching guarantees on two learners: the maximum likelihood estimator for the mean of a Gaussian, and the large margin classifier in 1D. For general lea…

Cited by 0SourcePDFScholar
2017

Learning Determinantal Point Processes with Moments and Cycles

ICML 2017poster

Determinantal Point Processes (DPPs) are a family of probabilistic models that have a repulsive behavior, and lend themselves naturally to many tasks in machine learning where returning a diverse set of objects is important. While there are fast algorithms for sampling, marginalization and condition…

Cited by 33SourcePDFScholar
2017

Near-linear time approximation algorithms for optimal transport via Sinkhorn iteration

NeurIPS 2017spotlight

Computing optimal transport distances such as the earth mover's distance is a fundamental problem in machine learning, statistics, and computer vision. Despite the recent introduction of several algorithms with good empirical performance, it is unknown whether general optimal transport distances can…

Cited by 754SourcePDFScholar