← Search

Lenka Zdeborova

30 accepted papers

2026

A Solvable High-Dimensional Model Where Nonlinear Autoencoders Learn Structure Invisible to PCA While Test Loss Misaligns With Generalization

ICML 2026poster

Many real-world datasets contain hidden structure that cannot be detected by simple linear correlations between input features. For example, latent factors may influence the data in a coordinated way, even though their effect is invisible to covariance-based methods such as PCA. In practice, nonline…

Cited by 0SourceScholar
2026

On the existence of consistent adversarial attacks in high-dimensional linear classification

ICML 2026spotlight

What fundamentally distinguishes an adversarial attack from a misclassification due to limited model expressivity or finite data? In this work, we investigate this question in the setting of high-dimensional binary classification, where statistical effects due to limited data availability play a cen…

Cited by 0SourceScholar
2026

Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling Laws

ICML 2026spotlight

Trained attention layers exhibit striking and reproducible spectral structure of the weights, including low-rank collapse, bulk deformation, and isolated spectral outliers, yet the origin of these phenomena and their implications for generalization remain poorly understood. We study empirical risk m…

Cited by 0SourceScholar
2025

Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention Networks

NeurIPS 2025poster

We study the dynamics of stochastic gradient descent (SGD) for a class of sequence models termed Sequence Single-Index (SSI) models, where the target depends on a single direction in input space applied to a sequence of tokens. This setting generalizes classical single-index models to the sequential…

Cited by 0SourceScholar
2025

Bayes optimal learning of attention-indexed models

NeurIPS 2025poster

We introduce the attention-indexed model (AIM), a theoretical framework for analyzing learning in deep attention layers. Inspired by multi-index models, AIM captures how token-level outputs emerge from layered bilinear interactions over high-dimensional embeddings. Unlike prior tractable attention m…

Cited by 0SourcecodeScholar
2025

Counting in Small Transformers: The Delicate Interplay between Attention and Feed-Forward Layers

ICML 2025poster

Next to scaling considerations, architectural design choices profoundly shape the solution space of transformers. In this work, we analyze the solutions simple transformer blocks implement when tackling the histogram task: counting items in sequences. Despite its simplicity, this task reveals a comp…

2025

Fundamental computational limits of weak learnability in high-dimensional multi-index models

AISTATS 2025poster

Multi-index models - functions which only depend on the covariates through a non-linear transformation of their projection on a subspace - are a useful benchmark for investigating feature learning with neural networks. This paper examines the theoretical boundaries of efficient learnability in this…

Cited by 0SourcecodeScholar
2025

Fundamental limits of learning in sequence multi-index models and deep attention networks: high-dimensional asymptotics and sharp thresholds

ICML 2025poster

In this manuscript, we study the learning of deep attention neural networks, defined as the composition of multiple self-attention layers, with tied and low-rank weights. We first establish a mapping of such models to sequence multi-index models, a generalization of the widely studied multi-index m…

2025

Learning with Restricted Boltzmann Machines: Asymptotics of AMP and GD in High Dimensions

NeurIPS 2025poster

The Restricted Boltzmann Machine (RBM) is one of the simplest generative neural networks capable of learning input distributions. Despite its simplicity, the analysis of its performance in learning from the training data is only well understood in cases that essentially reduce to singular value deco…

Cited by 0SourcecodeScholar
2025

The Computational Advantage of Depth in Learning High-Dimensional Hierarchical Targets

NeurIPS 2025spotlight

Understanding the advantages of deep neural networks trained by gradient descent (GD) compared to shallow models remains an open theoretical challenge. In this paper, we introduce a class of target functions (single and multi-index Gaussian hierarchical targets) that incorporate a hierarchy of laten…

Cited by 0SourceScholar
2025

The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks

NeurIPS 2025poster

We study the high-dimensional asymptotics of empirical risk minimization (ERM) in over-parametrized two-layer neural networks with quadratic activations trained on synthetic data. We derive sharp asymptotics for both training and test errors by mapping the $\ell_2$-regularized learning problem to a…

Cited by 0SourcecodeScholar
2024

A Phase Transition between Positional and Semantic Learning in a Solvable Model of Dot-Product Attention

NeurIPS 2024spotlight

Many empirical studies have provided evidence for the emergence of algorithmic mechanisms (abilities) in the learning of language models, that lead to qualitative improvements of the model capabilities. Yet, a theoretical characterization of how such mechanisms emerge remains elusive. In this paper,…

Cited by 13SourcePDFScholar
2024

Analysis of Learning a Flow-based Generative Model from Limited Sample Complexity

ICLR 2024poster

We study the problem of training a flow-based generative model, parametrized by a two-layer autoencoder, to sample from a high-dimensional Gaussian mixture. We provide a sharp end-to-end analysis of the problem. First, we provide a tight closed-form characterization of the learnt velocity field, whe…

2024

Asymptotics of feature learning in two-layer networks after one gradient-step

ICML 2024spotlight

In this manuscript, we investigate the problem of how two-layer neural networks learn features from data, and improve over the kernel regime, after being trained with a single gradient descent step. Leveraging the insight from (Ba et al., 2022), we model the trained network by a spiked Random Featur…

2024

Bayes-optimal learning of an extensive-width neural network from quadratically many samples

NeurIPS 2024poster

We consider the problem of learning a target function corresponding to a single hidden layer neural network, with a quadratic activation function after the first layer, and random weights. We consider the asymptotic limit where the input dimension and the network width are proportionally large. Rece…

2024

The Benefits of Reusing Batches for Gradient Descent in Two-Layer Networks: Breaking the Curse of Information and Leap Exponents

ICML 2024poster

We investigate the training dynamics of two-layer neural networks when learning multi-index target functions. We focus on multi-pass gradient descent (GD) that reuses the batches multiple times and show that it significantly changes the conclusion about which functions are learnable compared to sing…

2023

On double-descent in uncertainty quantification in overparametrized models

AISTATS 2023poster

Uncertainty quantification is a central challenge in reliable and trustworthy machine learning. Naive measures such as last-layer scores are well-known to yield overconfident estimates in the context of overparametrized neural networks. Several methods, ranging from temperature scaling to different…

2023

Universality laws for Gaussian mixtures in generalized linear models

NeurIPS 2023poster

A recent line of work in high-dimensional statistics working under the Gaussian mixture hypothesis has led to a number of results in the context of empirical risk minimization, Bayesian uncertainty quantification, separation of kernel methods and neural networks, ensembling and fluctuation of random…

Cited by 29SourcePDFScholar
2022

Multi-layer State Evolution Under Random Convolutional Design

NeurIPS 2022accept

Signal recovery under generative neural network priors has emerged as a promising direction in statistical inference and computational imaging. Theoretical analysis of reconstruction algorithms under generative priors is, however, challenging. For generative priors with fully connected layers and Ga…

2022

Phase diagram of Stochastic Gradient Descent in high-dimensional two-layer neural networks

NeurIPS 2022accept

Despite the non-convex optimization landscape, over-parametrized shallow networks are able to achieve global convergence under gradient descent. The picture can be radically different for narrow networks, which tend to get stuck in badly-generalizing local minima. Here we investigate the cross-over…

2022

Subspace clustering in high-dimensions: Phase transitions & Statistical-to-Computational gap

NeurIPS 2022accept

A simple model to study subspace clustering is the high-dimensional $k$-Gaussian mixture model where the cluster means are sparse vectors. Here we provide an exact asymptotic characterization of the statistically optimal reconstruction error in this model in the high-dimensional regime with extensiv…

2021

Classifying high-dimensional Gaussian mixtures: Where kernel methods fail and neural networks succeed

ICML 2021spotlight

A recent series of theoretical works showed that the dynamics of neural networks with a certain initialisation are well-captured by kernel methods. Concurrent empirical work demonstrated that kernel methods can come close to the performance of neural networks on some image classification tasks. Thes…

2021

Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy Regime

NeurIPS 2021poster

In this manuscript we consider Kernel Ridge Regression (KRR) under the Gaussian design. Exponents for the decay of the excess generalization error of KRR have been reported in various works under the assumption of power-law decay of eigenvalues of the features co-variance. These decays were, however…

Cited by 108SourcePDFScholar
2021

Learning Gaussian Mixtures with Generalized Linear Models: Precise Asymptotics in High-dimensions

NeurIPS 2021spotlight

Generalised linear models for multi-class classification problems are one of the fundamental building blocks of modern machine learning tasks. In this manuscript, we characterise the learning of a mixture of $K$ Gaussians with generic means and covariances via empirical risk minimisation (ERM) with…

Cited by 83SourcePDFScholar
2021

Learning curves of generic features maps for realistic datasets with a teacher-student model

NeurIPS 2021poster

Teacher-student models provide a framework in which the typical-case performance of high-dimensional supervised learning can be described in closed form. The assumptions of Gaussian i.i.d. input data underlying the canonical teacher-student model may, however, be perceived as too restrictive to capt…

2020

Generalisation error in learning with random features and the hidden manifold model

ICML 2020poster

We study generalised linear regression and classification for a synthetically generated dataset encompassing different problems of interest, such as learning with random features, neural networks in the lazy training regime, and the hidden manifold model. We consider the high-dimensional regime and…

Cited by 214SourcePDFScholar
2020

The Role of Regularization in Classification of High-dimensional Noisy Gaussian Mixture

ICML 2020poster

We consider a high-dimensional mixture of two Gaussians in the noisy regime where even an oracle knowing the centers of the clusters misclassifies a small but finite fraction of the points. We provide a rigorous analysis of the generalization error of regularized convex classifiers, including ridge,…

Cited by 111SourcePDFScholar
2019

Passed & Spurious: Descent Algorithms and Local Minima in Spiked Matrix-Tensor Models

ICML 2019oral

In this work we analyse quantitatively the interplay between the loss landscape and performance of descent algorithms in a prototypical inference problem, the spiked matrix-tensor model. We study a loss function that is the negative log-likelihood of the model. We analyse the number of local minima…

Cited by 67SourcePDFScholar