← Search

Florent Krzakala

62 accepted papers

2026

A Noise Sensitivity Exponent Controls Large Statistical-to-Computational Gaps in Single- and Multi-Index Models

ICML 2026spotlight

Understanding when learning is statistically possible yet computationally hard is a central challenge in high-dimensional statistics. In this work, we investigate this question in the context of single- and multi-index models, classes of functions widely studied as benchmarks to probe the ability of…

Cited by 0SourceScholar
2026

Efficient Learning of Compositional Targets with Hierarchical Spectral Methods

ICML 2026poster

Why depth yields a genuine computational advantage over shallow methods remains a central open question in learning theory. We study this question in a controlled high-dimensional Gaussian setting, focusing on compositional target functions. We analyze their learnability using an explicit three-laye…

Cited by 0SourceScholar
2026

Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime

ICLR 2026oral

Neural scaling laws underlie many of the recent advances in deep learning, yet their theoretical understanding remains largely confined to linear models. In this work, we present a systematic analysis of scaling laws for quadratic and diagonal neural networks in the feature learning regime. Leveragi…

Cited by 0SourcecodeScholar
2026

Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling Laws

ICML 2026spotlight

Trained attention layers exhibit striking and reproducible spectral structure of the weights, including low-rank collapse, bulk deformation, and isolated spectral outliers, yet the origin of these phenomena and their implications for generalization remain poorly understood. We study empirical risk m…

Cited by 0SourceScholar
2025

A High Dimensional Statistical Model for Adversarial Training: Geometry and Trade-Offs

AISTATS 2025poster

This work investigates adversarial training in the context of margin-based linear classifiers in the high-dimensional regime where the dimension $d$ and the number of data points $n$ diverge with a fixed ratio $\alpha = n / d$. We introduce a tractable mathematical model where the interplay betwee…

Cited by 0SourceScholar
2025

A Random Matrix Theory Perspective on the Spectrum of Learned Features and Asymptotic Generalization Capabilities

AISTATS 2025oral

A key property of neural networks is their capacity of adapting to data during training. Yet, our current mathematical understanding of feature learning and its relationship to generalization remain limited. In this work, we provide a random matrix analysis of how fully-connected two-layer neural ne…

Cited by 0SourceScholar
2025

Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention Networks

NeurIPS 2025poster

We study the dynamics of stochastic gradient descent (SGD) for a class of sequence models termed Sequence Single-Index (SSI) models, where the target depends on a single direction in input space applied to a sequence of tokens. This setting generalizes classical single-index models to the sequential…

Cited by 0SourceScholar
2025

Fundamental computational limits of weak learnability in high-dimensional multi-index models

AISTATS 2025poster

Multi-index models - functions which only depend on the covariates through a non-linear transformation of their projection on a subspace - are a useful benchmark for investigating feature learning with neural networks. This paper examines the theoretical boundaries of efficient learnability in this…

Cited by 0SourcecodeScholar
2025

Fundamental limits of learning in sequence multi-index models and deep attention networks: high-dimensional asymptotics and sharp thresholds

ICML 2025poster

In this manuscript, we study the learning of deep attention neural networks, defined as the composition of multiple self-attention layers, with tied and low-rank weights. We first establish a mapping of such models to sequence multi-index models, a generalization of the widely studied multi-index m…

2025

Learning with Restricted Boltzmann Machines: Asymptotics of AMP and GD in High Dimensions

NeurIPS 2025poster

The Restricted Boltzmann Machine (RBM) is one of the simplest generative neural networks capable of learning input distributions. Despite its simplicity, the analysis of its performance in learning from the training data is only well understood in cases that essentially reduce to singular value deco…

Cited by 0SourcecodeScholar
2025

Optimal Spectral Transitions in High-Dimensional Multi-Index Models

NeurIPS 2025poster

We consider the problem of how many samples from a Gaussian multi-index model are required to weakly reconstruct the relevant index subspace. Despite its increasing popularity as a testbed for investigating the computational complexity of neural networks, results beyond the single-index setting rema…

Cited by 0SourceScholar
2025

The Computational Advantage of Depth in Learning High-Dimensional Hierarchical Targets

NeurIPS 2025spotlight

Understanding the advantages of deep neural networks trained by gradient descent (GD) compared to shallow models remains an open theoretical challenge. In this paper, we introduce a class of target functions (single and multi-index Gaussian hierarchical targets) that incorporate a hierarchy of laten…

Cited by 0SourceScholar
2025

The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks

NeurIPS 2025poster

We study the high-dimensional asymptotics of empirical risk minimization (ERM) in over-parametrized two-layer neural networks with quadratic activations trained on synthetic data. We derive sharp asymptotics for both training and test errors by mapping the $\ell_2$-regularized learning problem to a…

Cited by 0SourcecodeScholar
2024

A Phase Transition between Positional and Semantic Learning in a Solvable Model of Dot-Product Attention

NeurIPS 2024spotlight

Many empirical studies have provided evidence for the emergence of algorithmic mechanisms (abilities) in the learning of language models, that lead to qualitative improvements of the model capabilities. Yet, a theoretical characterization of how such mechanisms emerge remains elusive. In this paper,…

Cited by 13SourcePDFScholar
2024

Analysis of Bootstrap and Subsampling in High-dimensional Regularized Regression

UAI 2024poster

We investigate popular resampling methods for estimating the uncertainty of statistical models, such as subsampling, bootstrap and the jackknife, and their performance in high-dimensional supervised regression tasks. We provide a tight asymptotic description of the biases and variances estimated by…

2024

Analysis of Learning a Flow-based Generative Model from Limited Sample Complexity

ICLR 2024poster

We study the problem of training a flow-based generative model, parametrized by a two-layer autoencoder, to sample from a high-dimensional Gaussian mixture. We provide a sharp end-to-end analysis of the problem. First, we provide a tight closed-form characterization of the learnt velocity field, whe…

2024

Asymptotic Characterisation of the Performance of Robust Linear Regression in the Presence of Outliers

AISTATS 2024poster

We study robust linear regression in high-dimension, when both the dimension $d$ and the number of data points $n$ diverge with a fixed ratio $\alpha=n/d$, and study a data model that includes outliers. We provide exact asymptotics for the performances of the empirical risk minimisation (ERM) using…

2024

Asymptotics of feature learning in two-layer networks after one gradient-step

ICML 2024spotlight

In this manuscript, we investigate the problem of how two-layer neural networks learn features from data, and improve over the kernel regime, after being trained with a single gradient descent step. Leveraging the insight from (Ba et al., 2022), we model the trained network by a spiked Random Featur…

2024

Bayes-optimal learning of an extensive-width neural network from quadratically many samples

NeurIPS 2024poster

We consider the problem of learning a target function corresponding to a single hidden layer neural network, with a quadratic activation function after the first layer, and random weights. We consider the asymptotic limit where the input dimension and the network width are proportionally large. Rece…

2024

Online Learning and Information Exponents: The Importance of Batch size & Time/Complexity Tradeoffs

ICML 2024poster

We study the impact of the batch size $n_b$ on the iteration time $T$ of training two-layer neural networks with one-pass stochastic gradient descent (SGD) on multi-index target functions of isotropic covariates. We characterize the optimal batch size minimizing the iteration time as a function of t…

Cited by 5SourcePDFScholar
2024

Spectral Phase Transition and Optimal PCA in Block-Structured Spiked Models

ICML 2024poster

We discuss the inhomogeneous Wigner spike model, a theoretical framework recently introduced to study structured noise in various learning scenarios, through the prism of random matrix theory, with a specific focus on its spectral properties. Our primary objective is to find an optimal spectral meth…

Cited by 1SourcePDFScholar
2024

The Benefits of Reusing Batches for Gradient Descent in Two-Layer Networks: Breaking the Curse of Information and Leap Exponents

ICML 2024poster

We investigate the training dynamics of two-layer neural networks when learning multi-index target functions. We focus on multi-pass gradient descent (GD) that reuses the batches multiple times and show that it significantly changes the conclusion about which functions are learnable compared to sing…

2023

Are Gaussian Data All You Need? The Extents and Limits of Universality in High-Dimensional Generalized Linear Estimation

ICML 2023poster

In this manuscript we consider the problem of generalized linear estimation on Gaussian mixture data with labels given by a single-index model. Our first result is a sharp asymptotic expression for the test and training errors in the high-dimensional regime. Motivated by the recent stream of results…

Cited by 34SourcePDFScholar
2023

Expectation consistency for calibration of neural networks

UAI 2023poster

Despite their incredible performance, it is well reported that deep neural networks tend to be overoptimistic about their prediction confidence. Finding effective and efficient calibration methods for neural networks is therefore an important endeavour towards better uncertainty quantification in de…

2023

On double-descent in uncertainty quantification in overparametrized models

AISTATS 2023poster

Uncertainty quantification is a central challenge in reliable and trustworthy machine learning. Naive measures such as last-layer scores are well-known to yield overconfident estimates in the context of overparametrized neural networks. Several methods, ranging from temperature scaling to different…

2023

Universality laws for Gaussian mixtures in generalized linear models

NeurIPS 2023poster

A recent line of work in high-dimensional statistics working under the Gaussian mixture hypothesis has led to a number of results in the context of empirical risk minimization, Bayesian uncertainty quantification, separation of kernel methods and neural networks, ensembling and fluctuation of random…

Cited by 29SourcePDFScholar
2022

Adversarial Robustness by Design Through Analog Computing And Synthetic Gradients

ICASSP 2022accepted

We propose a new defense mechanism against adversarial at-tacks inspired by an optical co-processor, providing robustness without compromising natural accuracy in both white-box and black-box settings. This hardware co-processor performs a nonlinear fixed random transformation, where the parameters…

Cited by 0SourceScholar
2022

Fluctuations, Bias, Variance & Ensemble of Learners: Exact Asymptotics for Convex Losses in High-Dimension

ICML 2022spotlight

From the sampling of data to the initialisation of parameters, randomness is ubiquitous in modern Machine Learning practice. Understanding the statistical fluctuations engendered by the different sources of randomness in prediction is therefore key to understanding robust generalisation. In this man…

Cited by 36SourcePDFScholar
2022

Multi-layer State Evolution Under Random Convolutional Design

NeurIPS 2022accept

Signal recovery under generative neural network priors has emerged as a promising direction in statistical inference and computational imaging. Theoretical analysis of reconstruction algorithms under generative priors is, however, challenging. For generative priors with fully connected layers and Ga…

2022

Phase diagram of Stochastic Gradient Descent in high-dimensional two-layer neural networks

NeurIPS 2022accept

Despite the non-convex optimization landscape, over-parametrized shallow networks are able to achieve global convergence under gradient descent. The picture can be radically different for narrow networks, which tend to get stuck in badly-generalizing local minima. Here we investigate the cross-over…

2022

Subspace clustering in high-dimensions: Phase transitions & Statistical-to-Computational gap

NeurIPS 2022accept

A simple model to study subspace clustering is the high-dimensional $k$-Gaussian mixture model where the cluster means are sparse vectors. Here we provide an exact asymptotic characterization of the statistically optimal reconstruction error in this model in the high-dimensional regime with extensiv…

2021

Classifying high-dimensional Gaussian mixtures: Where kernel methods fail and neural networks succeed

ICML 2021spotlight

A recent series of theoretical works showed that the dynamics of neural networks with a certain initialisation are well-captured by kernel methods. Concurrent empirical work demonstrated that kernel methods can come close to the performance of neural networks on some image classification tasks. Thes…

2021

Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy Regime

NeurIPS 2021poster

In this manuscript we consider Kernel Ridge Regression (KRR) under the Gaussian design. Exponents for the decay of the excess generalization error of KRR have been reported in various works under the assumption of power-law decay of eigenvalues of the features co-variance. These decays were, however…

Cited by 108SourcePDFScholar
2021

Learning Gaussian Mixtures with Generalized Linear Models: Precise Asymptotics in High-dimensions

NeurIPS 2021spotlight

Generalised linear models for multi-class classification problems are one of the fundamental building blocks of modern machine learning tasks. In this manuscript, we characterise the learning of a mixture of $K$ Gaussians with generic means and covariances via empirical risk minimisation (ERM) with…

Cited by 83SourcePDFScholar
2021

Learning curves of generic features maps for realistic datasets with a teacher-student model

NeurIPS 2021poster

Teacher-student models provide a framework in which the typical-case performance of high-dimensional supervised learning can be described in closed form. The assumptions of Gaussian i.i.d. input data underlying the canonical teacher-student model may, however, be perceived as too restrictive to capt…

2020

Complex Dynamics in Simple Neural Networks: Understanding Gradient Flow in Phase Retrieval

NeurIPS 2020poster

Despite the widespread use of gradient-based algorithms for optimising high-dimensional non-convex functions, understanding their ability of finding good minima instead of being trapped in spurious ones remains to a large extent an open problem. Here we focus on gradient flow dynamics for phase retr…

Cited by 36SourcePDFScholar
2020

Direct Feedback Alignment Scales to Modern Deep Learning Tasks and Architectures

NeurIPS 2020poster

Despite being the workhorse of deep learning, the backpropagation algorithm is no panacea. It enforces sequential layer updates, thus preventing efficient parallelization of the training process. Furthermore, its biological plausibility is being challenged. Alternative schemes have been devised; yet…

2020

Double Trouble in Double Descent: Bias and Variance(s) in the Lazy Regime

ICML 2020poster

Deep neural networks can achieve remarkable generalization performances while interpolating the training data. Rather than the U-curve emblematic of the bias-variance trade-off, their test error often follows a "double descent"—a mark of the beneficial role of overparametrization. In this work, we d…

2020

Dynamical mean-field theory for stochastic gradient descent in Gaussian mixture classification

NeurIPS 2020poster

We analyze in a closed form the learning dynamics of stochastic gradient descent (SGD) for a single layer neural network classifying a high-dimensional Gaussian mixture where each cluster is assigned one of two labels. This problem provides a prototype of a non-convex loss landscape with interpolati…

Cited by 107SourcePDFScholar
2020

Generalisation error in learning with random features and the hidden manifold model

ICML 2020poster

We study generalised linear regression and classification for a synthetically generated dataset encompassing different problems of interest, such as learning with random features, neural networks in the lazy training regime, and the hidden manifold model. We consider the high-dimensional regime and…

Cited by 214SourcePDFScholar
2020

Generalization error in high-dimensional perceptrons: Approaching Bayes error with convex optimization

NeurIPS 2020poster

We consider a commonly studied supervised classification of a synthetic dataset whose labels are generated by feeding a one-layer non-linear neural network with random iid inputs. We study the generalization performances of standard classifiers in the high-dimensional regime where $\alpha=\frac{n}{d…

Cited by 71SourcePDFScholar
2020

Kernel Computations from Large-Scale Random Features Obtained by Optical Processing Units

ICASSP 2020accepted

Approximating kernel functions with random features (RFs) has been a successful application of random projections for nonparametric estimation. However, performing random projections presents computational challenges for large-scale problems. Recently, a new optical hardware called Optical Processin…

Cited by 0SourceScholar
2020

Phase retrieval in high dimensions: Statistical and computational phase transitions

NeurIPS 2020poster

We consider the phase retrieval problem of reconstructing a $n$-dimensional real or complex signal $\mathbf{X}^\star$ from $m$ (possibly noisy) observations $Y_\mu = | \sum_{i=1}^n \Phi_{\mu i} X^{\star}_i/\sqrt{n}|$, for a large class of correlated real and complex random sensing matrices $\mathbf{…

2020

Reservoir Computing meets Recurrent Kernels and Structured Transforms

NeurIPS 2020oral

Reservoir Computing is a class of simple yet efficient Recurrent Neural Networks where internal weights are fixed at random and only a linear output layer is trained. In the large size limit, such random neural networks have a deep connection with kernel methods. Our contributions are threefold: a)…

2020

The Role of Regularization in Classification of High-dimensional Noisy Gaussian Mixture

ICML 2020poster

We consider a high-dimensional mixture of two Gaussians in the noisy regime where even an oracle knowing the centers of the clusters misclassifies a small but finite fraction of the points. We provide a rigorous analysis of the generalization error of regularized convex classifiers, including ridge,…

Cited by 111SourcePDFScholar
2019

Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup

NeurIPS 2019oral

Deep neural networks achieve stellar generalisation even when they have enough parameters to easily fit all their training data. We study this phenomenon by analysing the dynamics and the performance of over-parameterised two-layer neural networks in the teacher-student setup, where one network, the…

2019

Passed & Spurious: Descent Algorithms and Local Minima in Spiked Matrix-Tensor Models

ICML 2019oral

In this work we analyse quantitatively the interplay between the loss landscape and performance of descent algorithms in a prototypical inference problem, the spiked matrix-tensor model. We study a loss function that is the negative log-likelihood of the model. We analyse the number of local minima…

Cited by 67SourcePDFScholar
2019

Spectral Method for Multiplexed Phase Retrieval and Application in Optical Imaging in Complex Media

ICASSP 2019accepted

We introduce a generalized version of phase retrieval called multiplexed phase retrieval. We want to recover the phase of amplitude-only measurements from linear combinations of them. This corresponds to the case in which multiple incoherent sources are sampled jointly, and one would like to recover…

Cited by 0SourceScholar
2019

The spiked matrix model with generative priors

NeurIPS 2019poster

Using a low-dimensional parametrization of signals is a generic and powerful way to enhance performance in signal processing and statistical inference. A very popular and widely explored type of dimensionality reduction is sparsity; another type is generative modelling of signal distributions. Gener…

2019

Who is Afraid of Big Bad Minima? Analysis of gradient-flow in spiked matrix-tensor models

NeurIPS 2019spotlight

Gradient-based algorithms are effective for many machine learning tasks, but despite ample recent effort and some progress, it often remains unclear why they work in practice in optimising high-dimensional non-convex functions and why they find good minima instead of being trapped in spurious ones.H…

2018

Entropy and mutual information in models of deep neural networks

NeurIPS 2018spotlight

We examine a class of stochastic deep learning models with a tractable method to compute information-theoretic quantities. Our contributions are three-fold: (i) We show how entropies and mutual informations can be derived from heuristic statistical physics methods, under the assumption that weight m…

Cited by 232SourcePDFScholar
2018

The committee machine: Computational to statistical gaps in learning a two-layers neural network

NeurIPS 2018spotlight

Heuristic tools from statistical physics have been used in the past to compute the optimal learning and generalization errors in the teacher-student scenario in multi- layer neural networks. In this contribution, we provide a rigorous justification of these approaches for a two-layers neural network…

2016

Intensity-only optical compressive imaging using a multiply scattering material and a double phase retrieval approach

ICASSP 2016accepted

In this paper, the problem of compressive imaging is addressed using natural randomization by means of a multiply scattering medium. To utilize the medium in this way, its corresponding transmission matrix must be estimated. For calibration purposes, we use a digital micromirror device (DMD) as a si…

Cited by 0SourceScholar
2016

Mutual information for symmetric rank-one matrix estimation: A proof of the replica formula

NeurIPS 2016poster

Factorizing low-rank matrices has many applications in machine learning and statistics. For probabilistic models in the Bayes optimal setting, a general expression for the mutual information has been proposed using heuristic statistical physics computations, and proven in few specific cases. Here, w…

Cited by 217SourcePDFScholar
2016

Random projections through multiple optical scattering: Approximating Kernels at the speed of light

ICASSP 2016accepted

Random projections have proven extremely useful in many signal processing and machine learning applications. However, they often require either to store a very large random matrix, or to use a different, structured matrix to reduce the computational and memory costs. Here, we overcome this difficult…

Cited by 0SourceScholar
2015

Adaptive damping and mean removal for the generalized approximate message passing algorithm

ICASSP 2015accepted

The generalized approximate message passing (GAMP) algorithm is an efficient method of MAP or approximate-MMSE estimation of x observed from a noisy version of the transform coefficients z = Ax. In fact, for large zero-mean i.i.d sub-Gaussian A, GAMP is characterized by a state evolution whose fixed…

Cited by 0SourceScholar
2015

Matrix Completion from Fewer Entries: Spectral Detectability and Rank Estimation

NeurIPS 2015poster

The completion of low rank matrices from few entries is a task with many practical applications. We consider here two aspects of this problem: detectability, i.e. the ability to estimate the rank $r$ reliably from the fewest possible random entries, and performance in achieving small reconstruction…

Cited by 22SourcePDFScholar
2015

Swept Approximate Message Passing for Sparse Estimation

ICML 2015poster

Approximate Message Passing (AMP) has been shown to be a superior method for inference problems, such as the recovery of signals from sets of noisy, lower-dimensionality measurements, both in terms of reconstruction accuracy and in computational efficiency. However, AMP suffers from serious converge…

Cited by 52SourcePDFScholar
2015

Training Restricted Boltzmann Machine via the Thouless-Anderson-Palmer free energy

NeurIPS 2015poster

Restricted Boltzmann machines are undirected neural networks which have been shown tobe effective in many applications, including serving as initializations fortraining deep multi-layer neural networks. One of the main reasons for their success is theexistence of efficient and practical stochastic a…