← Search

Gerard Ben Arous

5 accepted papers

2025

Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws

NeurIPS 2025poster

We study the optimization and sample complexity of gradient-based training of a two-layer neural network with quadratic activation function in the high-dimensional regime, where the data is generated as $y \propto \sum_{j=1}^{r}\lambda_j \sigma\left(\langle \boldsymbol{\theta_j}, \boldsymbol{x}\rang…

Cited by 0SourceScholar
2024

High-dimensional SGD aligns with emerging outlier eigenspaces

ICLR 2024spotlight

We rigorously study the joint evolution of training dynamics via stochastic gradient descent (SGD) and the spectra of empirical Hessian and gradient matrices. We prove that in two canonical classification tasks for multi-class high-dimensional mixtures and either 1 or 2-layer neural networks, the SG…

Cited by 11SourcePDFScholar
2022

High-dimensional limit theorems for SGD: Effective dynamics and critical scaling

NeurIPS 2022accept

We study the scaling limits of stochastic gradient descent (SGD) with constant step-size in the high-dimensional regime. We prove limit theorems for the trajectories of summary statistics (i.e., finite-dimensional functions) of SGD as the dimension goes to infinity. Our approach allows one to choose…

Cited by 87SourcePDFScholar
2018

Comparing Dynamics: Deep Neural Networks versus Glassy Systems

ICML 2018oral

We analyze numerically the training dynamics of deep neural networks (DNN) by using methods developed in statistical physics of glassy systems. The two main issues we address are the complexity of the loss-landscape and of the dynamics within it, and to what extent DNNs share similarities with glass…

Cited by 138SourcePDFScholar
2015

The Loss Surfaces of Multilayer Networks

AISTATS 2015poster

We study the connection between the highly non-convex loss function of a simple model of the fully-connected feed-forward neural network and the Hamiltonian of the spherical spin-glass model under the assumptions of: i) variable independence, ii) redundancy in network parametrization, and iii) unifo…

Cited by 1741SourcePDFScholar