← Search

Giulio Biroli

15 accepted papers

2026

Overshoot and Shrinkage in Classifier-Free Guidance: From Theory to Practice

ICLR 2026poster

Classifier-Free Guidance (CFG) is widely used in diffusion and flow-based generative models for high-quality conditional generation, yet its theoretical properties remain incompletely understood. By connecting CFG to the high-dimensional framework of diffusion regimes, we show that in sufficiently h…

Cited by 0SourceScholar
2025

A Differentiable Rank-Based Objective for Better Feature Learning

ICLR 2025poster

In this paper, we leverage existing statistical methods to better understand feature learning from data. We tackle this by modifying the model-free variable selection method, Feature Ordering by Conditional Independence (FOCI), which is introduced in Azadkia & Chatterjee (2021). While FOCI is based…

Cited by 0SourcePDFScholar
2025

Why Diffusion Models Don’t Memorize: The Role of Implicit Dynamical Regularization in Training

NeurIPS 2025oral

Diffusion models have achieved remarkable success across a wide range of generative tasks. A key challenge is understanding the mechanisms that prevent their memorization of training data and allow generalization. In this work, we investigate the role of the training dynamics in the transition from…

Cited by 0SourceScholar
2024

Cascade of phase transitions in the training of energy-based models

NeurIPS 2024poster

In this paper, we investigate the feature encoding process in a prototypical energy-based generative model, the Restricted Boltzmann Machine (RBM). We start with an analytical investigation using simplified architectures and data structures, and end with numerical analysis of real trainings on real…

Cited by 3SourcePDFScholar
2024

On the Impact of Overparameterization on the Training of a Shallow Neural Network in High Dimensions

AISTATS 2024poster

We study the training dynamics of a shallow neural network with quadratic activation functions and quadratic cost in a teacher-student setup. In line with previous works on the same neural architecture, the optimization is performed following the gradient flow on the population risk, where the avera…

Cited by 11SourcePDFScholar
2022

Neural Network Pruning Denoises the Features and Makes Local Connectivity Emerge in Visual Tasks

ICML 2022spotlight

Pruning methods can considerably reduce the size of artificial neural networks without harming their performance and in some cases they can even uncover sub-networks that, when trained in isolation, match or surpass the test accuracy of their dense counterparts. Here, we characterize the inductive b…

Cited by 12SourcePDFScholar
2021

ConViT: Improving Vision Transformers with Soft Convolutional Inductive Biases

ICML 2021spotlight

Convolutional architectures have proven extremely successful for vision tasks. Their hard inductive biases enable sample-efficient learning, but come at the cost of a potentially lower performance ceiling. Vision Transformers (ViTs) rely on more flexible self-attention layers, and have recently outp…

2021

On the interplay between data structure and loss function in classification problems

NeurIPS 2021poster

One of the central features of modern machine learning models, including deep neural networks, is their generalization ability on structured data in the over-parametrized regime. In this work, we consider an analytically solvable setup to investigate how properties of data impact learning in classi…

2020

An analytic theory of shallow networks dynamics for hinge loss classification

NeurIPS 2020poster

Neural networks have been shown to perform incredibly well in classification tasks over structured high-dimensional datasets. However, the learning dynamics of such networks is still poorly understood. In this paper we study in detail the training dynamics of a simple type of neural network: a singl…

2020

Complex Dynamics in Simple Neural Networks: Understanding Gradient Flow in Phase Retrieval

NeurIPS 2020poster

Despite the widespread use of gradient-based algorithms for optimising high-dimensional non-convex functions, understanding their ability of finding good minima instead of being trapped in spurious ones remains to a large extent an open problem. Here we focus on gradient flow dynamics for phase retr…

Cited by 36SourcePDFScholar
2020

Double Trouble in Double Descent: Bias and Variance(s) in the Lazy Regime

ICML 2020poster

Deep neural networks can achieve remarkable generalization performances while interpolating the training data. Rather than the U-curve emblematic of the bias-variance trade-off, their test error often follows a "double descent"—a mark of the beneficial role of overparametrization. In this work, we d…

2020

Triple descent and the two kinds of overfitting: where & why do they appear?

NeurIPS 2020spotlight

A recent line of research has highlighted the existence of a ``double descent'' phenomenon in deep learning, whereby increasing the number of training examples N causes the generalization error of neural networks to peak when N is of the same order as the number of parameters P. In earlier works, a…

2019

Finding the Needle in the Haystack with Convolutions: on the benefits of architectural bias

NeurIPS 2019poster

Despite the phenomenal success of deep neural networks in a broad range of learning tasks, there is a lack of theory to understand the way they work. In particular, Convolutional Neural Networks (CNNs) are known to perform much better than Fully-Connected Networks (FCNs) on spatially structured data…

2019

Who is Afraid of Big Bad Minima? Analysis of gradient-flow in spiked matrix-tensor models

NeurIPS 2019spotlight

Gradient-based algorithms are effective for many machine learning tasks, but despite ample recent effort and some progress, it often remains unclear why they work in practice in optimising high-dimensional non-convex functions and why they find good minima instead of being trapped in spurious ones.H…

2018

Comparing Dynamics: Deep Neural Networks versus Glassy Systems

ICML 2018oral

We analyze numerically the training dynamics of deep neural networks (DNN) by using methods developed in statistical physics of glassy systems. The two main issues we address are the complexity of the loss-landscape and of the dynamics within it, and to what extent DNNs share similarities with glass…

Cited by 138SourcePDFScholar