← Search

Ludovic STEPHAN

5 accepted papers

2025

Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention Networks

NeurIPS 2025poster

We study the dynamics of stochastic gradient descent (SGD) for a class of sequence models termed Sequence Single-Index (SSI) models, where the target depends on a single direction in input space applied to a sequence of tokens. This setting generalizes classical single-index models to the sequential…

Cited by 0SourceScholar
2024

Online Learning and Information Exponents: The Importance of Batch size & Time/Complexity Tradeoffs

ICML 2024poster

We study the impact of the batch size $n_b$ on the iteration time $T$ of training two-layer neural networks with one-pass stochastic gradient descent (SGD) on multi-index target functions of isotropic covariates. We characterize the optimal batch size minimizing the iteration time as a function of t…

Cited by 5SourcePDFScholar
2023

Are Gaussian Data All You Need? The Extents and Limits of Universality in High-Dimensional Generalized Linear Estimation

ICML 2023poster

In this manuscript we consider the problem of generalized linear estimation on Gaussian mixture data with labels given by a single-index model. Our first result is a sharp asymptotic expression for the test and training errors in the high-dimensional regime. Motivated by the recent stream of results…

Cited by 34SourcePDFScholar
2023

Universality laws for Gaussian mixtures in generalized linear models

NeurIPS 2023poster

A recent line of work in high-dimensional statistics working under the Gaussian mixture hypothesis has led to a number of results in the context of empirical risk minimization, Bayesian uncertainty quantification, separation of kernel methods and neural networks, ensembling and fluctuation of random…

Cited by 29SourcePDFScholar
2022

Phase diagram of Stochastic Gradient Descent in high-dimensional two-layer neural networks

NeurIPS 2022accept

Despite the non-convex optimization landscape, over-parametrized shallow networks are able to achieve global convergence under gradient descent. The picture can be radically different for narrow networks, which tend to get stuck in badly-generalizing local minima. Here we investigate the cross-over…