← Search

Arthur Jacot

16 accepted papers

2026

Saddle-To-Saddle Dynamics in Deep ReLU Networks: Low-Rank Bias in the First Saddle Escape

ICLR 2026poster

When a deep ReLU network is initialized with small weights, gradient descent (GD) is at first dominated by the saddle at the origin in parameter space. We study the so-called escape directions along which GD leaves the origin, which play a similar role as the eigenvectors of the Hessian for strict s…

Cited by 0SourceScholar
2025

How DNNs break the Curse of Dimensionality: Compositionality and Symmetry Learning

ICLR 2025poster

We show that deep neural networks (DNNs) can efficiently learn any composition of functions with bounded $F_{1}$-norm, which allows DNNs to break the curse of dimensionality in ways that shallow networks cannot. More specifically, we derive a generalization bound that combines a covering number argu…

2025

Shallow diffusion networks provably learn hidden low-dimensional structure

ICLR 2025poster

Diffusion-based generative models provide a powerful framework for learning to sample from a complex target distribution. The remarkable empirical success of these models applied to high-dimensional signals, including images and video, stands in stark contrast to classical results highlighting the c…

Cited by 4SourcePDFScholar
2025

Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse

ICLR 2025oral

Deep neural networks (DNNs) at convergence consistently represent the training data in the last layer via a geometric structure referred to as neural collapse. This empirical evidence has spurred a line of theoretical research aimed at proving the emergence of neural collapse, mostly focusing on the…

Cited by 2SourcePDFScholar
2024

Implicit bias of SGD in $L_2$-regularized linear DNNs: One-way jumps from high to low rank

ICLR 2024spotlight

The $L_{2}$-regularized loss of Deep Linear Networks (DLNs) with more than one hidden layers has multiple local minima, corresponding to matrices with different ranks. In tasks such as matrix completion, the goal is to converge to the local minimum with the smallest rank that still fits the training…

Cited by 20SourcePDFScholar
2024

Mixed Dynamics In Linear Networks: Unifying the Lazy and Active Regimes

NeurIPS 2024poster

The training dynamics of linear networks are well studied in two distinct setups: the lazy regime and balanced/active regime, depending on the initialization and width of the network. We provide a surprisingly simple unifying formula for the evolution of the learned matrix that contains as special c…

Cited by 7SourcePDFScholar
2022

Feature Learning in $L_2$-regularized DNNs: Attraction/Repulsion and Sparsity

NeurIPS 2022accept

We study the loss surface of DNNs with $L_{2}$ regularization. We show that the loss in terms of the parameters can be reformulated into a loss in terms of the layerwise activations $Z_{\ell}$ of the training set. This reformulation reveals the dynamics behind feature learning: each hidden represent…

Cited by 22SourcePDFScholar
2021

DNN-based Topology Optimisation: Spatial Invariance and Neural Tangent Kernel

NeurIPS 2021poster

We study the Solid Isotropic Material Penalization (SIMP) method with a density field generated by a fully-connected neural network, taking the coordinates as inputs. In the large width limit, we show that the use of DNNs leads to a filtering effect similar to traditional filtering techniques for SI…

2021

Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances

ICML 2021spotlight

We study how permutation symmetries in overparameterized multi-layer neural networks generate ‘symmetry-induced’ critical points. Assuming a network with $ L $ layers of minimal widths $ r_1^*, \ldots, r_{L-1}^* $ reaches a zero-loss minimum at $ r_1^*! \cdots r_{L-1}^*! $ isolated points that are p…

2020

Implicit Regularization of Random Feature Models

ICML 2020poster

Random Features (RF) models are used as efficient parametric approximations of kernel methods. We investigate, by means of random matrix theory, the connection between Gaussian RF models and Kernel Ridge Regression (KRR). For a Gaussian RF model with $P$ features, $N$ data points, and a ridge $\lamb…

Cited by 108SourcePDFScholar
2020

Kernel Alignment Risk Estimator: Risk Prediction from Training Data

NeurIPS 2020poster

We study the risk (i.e. generalization error) of Kernel Ridge Regression (KRR) for a kernel $K$ with ridge $\lambda>0$ and i.i.d. observations. For this, we introduce two objects: the Signal Capture Threshold (SCT) and the Kernel Alignment Risk Estimator (KARE). The SCT $\vartheta_{K,\lambda}$ is a…

Cited by 72SourcePDFScholar
2018

Neural Tangent Kernel: Convergence and Generalization in Neural Networks

NeurIPS 2018spotlight

At initialization, artificial neural networks (ANNs) are equivalent to Gaussian processes in the infinite-width limit, thus connecting them to kernel methods. We prove that the evolution of an ANN during training can also be described by a kernel: during gradient descent on the parameters of an ANN,…

Cited by 4205SourcePDFScholar