← Search

Loucas Pillaud-Vivien

12 accepted papers

2025

Convergence of the Gradient Flow for Shallow ReLU Networks on Weakly Interacting Data

NeurIPS 2025poster

We analyse the convergence of one-hidden-layer ReLU networks trained by gradient flow on $n$ data points. Our main contribution leverages the high dimensionality of the ambient space, which implies low correlation of the input samples, to demonstrate that a network with width of order $\log(n)$ neur…

Cited by 0SourceScholar
2024

Batch and match: black-box variational inference with a score-based divergence

ICML 2024spotlight

Most leading implementations of black-box variational inference (BBVI) are based on optimizing a stochastic evidence lower bound (ELBO). But such approaches to BBVI often converge slowly due to the high variance of their gradient estimates and their sensitivity to hyperparameters. In this work, we p…

Cited by 7SourcePDFScholar
2023

On the spectral bias of two-layer linear networks

NeurIPS 2023poster

This paper studies the behaviour of two-layer fully connected networks with linear activations trained with gradient flow on the square loss. We show how the optimization process carries an implicit bias on the parameters that depends on the scale of its initialization. The main result of the paper…

Cited by 15SourcePDFScholar
2023

SGD with Large Step Sizes Learns Sparse Features

ICML 2023poster

We showcase important features of the dynamics of the Stochastic Gradient Descent (SGD) in the training of neural networks. We present empirical observations that commonly used large step sizes (i) may lead the iterates to jump from one side of a valley to the other causing *loss stabilization*, and…

2022

Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs

NeurIPS 2022accept

The training of neural networks by gradient descent methods is a cornerstone of the deep learning revolution. Yet, despite some recent progress, a complete theory explaining its success is still missing. This article presents, for orthogonal input vectors, a precise description of the gradient flow…

2021

Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of Stochasticity

NeurIPS 2021poster

Understanding the implicit bias of training algorithms is of crucial importance in order to explain the success of overparametrised neural networks. In this paper, we study the dynamics of stochastic gradient descent over diagonal linear networks through its continuous time version, namely stochasti…

Cited by 127SourcePDFScholar
2021

Last iterate convergence of SGD for Least-Squares in the Interpolation regime.

NeurIPS 2021poster

Motivated by the recent successes of neural networks that have the ability to fit the data perfectly \emph{and} generalize well, we study the noiseless model in the fundamental least-squares setup. We assume that an optimum predictor perfectly fits the inputs and outputs $\langle \theta_* , \phi(X)…

Cited by 45SourcePDFScholar
2021

Overcoming the curse of dimensionality with Laplacian regularization in semi-supervised learning

NeurIPS 2021poster

As annotations of data can be scarce in large-scale practical problems, leveraging unlabelled examples is one of the most important aspects of machine learning. This is the aim of semi-supervised learning. To benefit from the access to unlabelled data, it is natural to diffuse smoothly knowledge of…

2020

Statistical Estimation of the Poincaré constant and Application to Sampling Multimodal Distributions

AISTATS 2020poster

Poincaré inequalities are ubiquitous in probability and analysis and have various applications in statistics (concentration of measure, rate of convergence of Markov chains). The Poincaré constant, for which the inequality is tight, is related to the typical convergence rate of diffusions to their e…

Cited by 22SourcePDFScholar
2018

Statistical Optimality of Stochastic Gradient Descent on Hard Learning Problems through Multiple Passes

NeurIPS 2018poster

We consider stochastic gradient descent (SGD) for least-squares regression with potentially several passes over the data. While several passes have been widely reported to perform practically better in terms of predictive performance on unseen data, the existing theoretical analysis of SGD suggests…

Cited by 128SourcePDFScholar