← Search

Lénaïc Chizat

11 accepted papers

2024

Deep linear networks for regression are implicitly regularized towards flat minima

NeurIPS 2024poster

The largest eigenvalue of the Hessian, or sharpness, of neural networks is a key quantity to understand their optimization dynamics. In this paper, we study the sharpness of deep linear networks for univariate regression. Minimizers can have arbitrarily large sharpness, but not an arbitrarily small…

2024

Mean-Field Langevin Dynamics for Signed Measures via a Bilevel Approach

NeurIPS 2024spotlight

Mean-field Langevin dynamics (MLFD) is a class of interacting particle methods that tackle convex optimization over probability measures on a manifold, which are scalable, versatile, and enjoy computational guarantees. However, some important problems -- such as risk minimization for infinite width…

2024

The Feature Speed Formula: a flexible approach to scale hyper-parameters of deep neural networks

NeurIPS 2024poster

Deep learning succeeds by doing hierarchical feature learning, yet tuning hyper-parameters (HP) such as initialization scales, learning rates etc., only give indirect control over this behavior. In this paper, we introduce a key notion to predict and control feature learning: the angle $\theta_\ell$…

Cited by 1SourcePDFScholar
2023

Local Convergence of Gradient Methods for Min-Max Games: Partial Curvature Generically Suffices

NeurIPS 2023poster

We study the convergence to local Nash equilibria of gradient methods for two-player zero-sum differentiable games. It is well-known that, in the continuous-time setting, such dynamics converge locally when $S \succ 0$ and may diverge when $S=0$, where $S\succeq 0$ is the symmetric part of the Jacob…

Cited by 3SourcePDFScholar
2022

Trajectory Inference via Mean-field Langevin in Path Space

NeurIPS 2022accept

Trajectory inference aims at recovering the dynamics of a population from snapshots of its temporal marginals. To solve this task, a min-entropy estimator relative to the Wiener measure in path space was introduced in [Lavenant et al., 2021], and shown to consistently recover the dynamics of a large…

2020

Faster Wasserstein Distance Estimation with the Sinkhorn Divergence

NeurIPS 2020poster

The squared Wasserstein distance is a natural quantity to compare probability distributions in a non-parametric setting. This quantity is usually estimated with the plug-in estimator, defined via a discrete optimal transport problem which can be solved to $\epsilon$-accuracy by adding an entropic re…

Cited by 211SourcePDFScholar
2020

Statistical and Topological Properties of Sliced Probability Divergences

NeurIPS 2020spotlight

The idea of slicing divergences has been proven to be successful when comparing two probability measures in various machine learning applications including generative modeling, and consists in computing the expected value of a `base divergence' between \emph{one-dimensional random projections} of th…

2019

Sample Complexity of Sinkhorn Divergences

AISTATS 2019poster

Optimal transport (OT) and maximum mean discrepancies (MMD) are now routinely used in machine learning to compare probability measures. We focus in this paper on Sinkhorn divergences (SDs), a regularized variant of OT distances which can interpolate, depending on the regularization strength $\varep…

Cited by 362SourcePDFScholar
2018

On the Global Convergence of Gradient Descent for Over-parameterized Models using Optimal Transport

NeurIPS 2018poster

Many tasks in machine learning and signal processing can be solved by minimizing a convex function of a measure. This includes sparse spikes deconvolution or training a neural network with a single hidden layer. For these problems, we study a simple minimization method: the unknown measure is discre…

Cited by 961SourcePDFScholar