← Search

Pierfrancesco Urbani

6 accepted papers

2025

Dynamical Decoupling of Generalization and Overfitting in Large Two-Layer Networks

NeurIPS 2025oral

Understanding the inductive bias and generalization properties of large overparametrized machine learning models requires to characterize the dynamics of the training algorithm. We study the learning dynamics of large two-layer neural networks via dynamical mean field theory, a well established tec…

Cited by 0SourceScholar
2021

Analytical Study of Momentum-Based Acceleration Methods in Paradigmatic High-Dimensional Non-Convex Problems

NeurIPS 2021poster

The optimization step in many machine learning problems rarely relies on vanilla gradient descent but it is common practice to use momentum-based accelerated methods. Despite these algorithms being widely applied to arbitrary loss functions, their behaviour in generically non-convex, high dimensiona…

Cited by 14SourcePDFScholar
2020

Complex Dynamics in Simple Neural Networks: Understanding Gradient Flow in Phase Retrieval

NeurIPS 2020poster

Despite the widespread use of gradient-based algorithms for optimising high-dimensional non-convex functions, understanding their ability of finding good minima instead of being trapped in spurious ones remains to a large extent an open problem. Here we focus on gradient flow dynamics for phase retr…

Cited by 36SourcePDFScholar
2020

Dynamical mean-field theory for stochastic gradient descent in Gaussian mixture classification

NeurIPS 2020poster

We analyze in a closed form the learning dynamics of stochastic gradient descent (SGD) for a single layer neural network classifying a high-dimensional Gaussian mixture where each cluster is assigned one of two labels. This problem provides a prototype of a non-convex loss landscape with interpolati…

Cited by 107SourcePDFScholar
2020

The Role of Regularization in Classification of High-dimensional Noisy Gaussian Mixture

ICML 2020poster

We consider a high-dimensional mixture of two Gaussians in the noisy regime where even an oracle knowing the centers of the clusters misclassifies a small but finite fraction of the points. We provide a rigorous analysis of the generalization error of regularized convex classifiers, including ridge,…

Cited by 111SourcePDFScholar
2019

Passed & Spurious: Descent Algorithms and Local Minima in Spiked Matrix-Tensor Models

ICML 2019oral

In this work we analyse quantitatively the interplay between the loss landscape and performance of descent algorithms in a prototypical inference problem, the spiked matrix-tensor model. We study a loss function that is the negative log-likelihood of the model. We analyse the number of local minima…

Cited by 67SourcePDFScholar