← Search

Lenka Zdeborová

19 accepted papers

2026

Dataset Distillation for Memorized Data: Soft Labels can Leak Held-Out Teacher Knowledge

ICLR 2026poster

Dataset distillation aims to compress training data into fewer examples via a teacher, from which a student can learn effectively. While its success is often attributed to structure in the data, modern neural networks also memorize specific facts, but if and how such memorized information can be tra…

Cited by 0SourcecodeScholar
2026

Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime

ICLR 2026oral

Neural scaling laws underlie many of the recent advances in deep learning, yet their theoretical understanding remains largely confined to linear models. In this work, we present a systematic analysis of scaling laws for quadratic and diagonal neural networks in the feature learning regime. Leveragi…

Cited by 16SourcecodeScholar
2026

Statistical Advantage of Softmax Attention: Insights from Single-Location Regression

ICLR 2026poster

Large language models rely on attention mechanisms with a softmax activation. Yet the dominance of softmax over alternatives (e.g., component-wise or linear) remains poorly understood, and many theoretical works have focused on the easier-to-analyze linearized attention. In this work, we address thi…

Cited by 0SourcecodeScholar
2024

Analysis of Bootstrap and Subsampling in High-dimensional Regularized Regression

UAI 2024poster

We investigate popular resampling methods for estimating the uncertainty of statistical models, such as subsampling, bootstrap and the jackknife, and their performance in high-dimensional supervised regression tasks. We provide a tight asymptotic description of the biases and variances estimated by…

2023

Expectation consistency for calibration of neural networks

UAI 2023poster

Despite their incredible performance, it is well reported that deep neural networks tend to be overoptimistic about their prediction confidence. Finding effective and efficient calibration methods for neural networks is therefore an important endeavour towards better uncertainty quantification in de…

2020

Complex Dynamics in Simple Neural Networks: Understanding Gradient Flow in Phase Retrieval

NeurIPS 2020poster

Despite the widespread use of gradient-based algorithms for optimising high-dimensional non-convex functions, understanding their ability of finding good minima instead of being trapped in spurious ones remains to a large extent an open problem. Here we focus on gradient flow dynamics for phase retr…

Cited by 36SourcePDFScholar
2020

Dynamical mean-field theory for stochastic gradient descent in Gaussian mixture classification

NeurIPS 2020poster

We analyze in a closed form the learning dynamics of stochastic gradient descent (SGD) for a single layer neural network classifying a high-dimensional Gaussian mixture where each cluster is assigned one of two labels. This problem provides a prototype of a non-convex loss landscape with interpolati…

Cited by 107SourcePDFScholar
2020

Generalization error in high-dimensional perceptrons: Approaching Bayes error with convex optimization

NeurIPS 2020poster

We consider a commonly studied supervised classification of a synthetic dataset whose labels are generated by feeding a one-layer non-linear neural network with random iid inputs. We study the generalization performances of standard classifiers in the high-dimensional regime where $\alpha=\frac{n}{d…

Cited by 71SourcePDFScholar
2020

Phase retrieval in high dimensions: Statistical and computational phase transitions

NeurIPS 2020poster

We consider the phase retrieval problem of reconstructing a $n$-dimensional real or complex signal $\mathbf{X}^\star$ from $m$ (possibly noisy) observations $Y_\mu = | \sum_{i=1}^n \Phi_{\mu i} X^{\star}_i/\sqrt{n}|$, for a large class of correlated real and complex random sensing matrices $\mathbf{…

2019

Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup

NeurIPS 2019oral

Deep neural networks achieve stellar generalisation even when they have enough parameters to easily fit all their training data. We study this phenomenon by analysing the dynamics and the performance of over-parameterised two-layer neural networks in the teacher-student setup, where one network, the…

2019

The spiked matrix model with generative priors

NeurIPS 2019poster

Using a low-dimensional parametrization of signals is a generic and powerful way to enhance performance in signal processing and statistical inference. A very popular and widely explored type of dimensionality reduction is sparsity; another type is generative modelling of signal distributions. Gener…

2019

Who is Afraid of Big Bad Minima? Analysis of gradient-flow in spiked matrix-tensor models

NeurIPS 2019spotlight

Gradient-based algorithms are effective for many machine learning tasks, but despite ample recent effort and some progress, it often remains unclear why they work in practice in optimising high-dimensional non-convex functions and why they find good minima instead of being trapped in spurious ones.H…

2018

Entropy and mutual information in models of deep neural networks

NeurIPS 2018spotlight

We examine a class of stochastic deep learning models with a tractable method to compute information-theoretic quantities. Our contributions are three-fold: (i) We show how entropies and mutual informations can be derived from heuristic statistical physics methods, under the assumption that weight m…

Cited by 232SourcePDFScholar
2018

The committee machine: Computational to statistical gaps in learning a two-layers neural network

NeurIPS 2018spotlight

Heuristic tools from statistical physics have been used in the past to compute the optimal learning and generalization errors in the teacher-student scenario in multi- layer neural networks. In this contribution, we provide a rigorous justification of these approaches for a two-layers neural network…

2016

Mutual information for symmetric rank-one matrix estimation: A proof of the replica formula

NeurIPS 2016poster

Factorizing low-rank matrices has many applications in machine learning and statistics. For probabilistic models in the Bayes optimal setting, a general expression for the mutual information has been proposed using heuristic statistical physics computations, and proven in few specific cases. Here, w…

Cited by 217SourcePDFScholar
2015

Adaptive damping and mean removal for the generalized approximate message passing algorithm

ICASSP 2015accepted

The generalized approximate message passing (GAMP) algorithm is an efficient method of MAP or approximate-MMSE estimation of x observed from a noisy version of the transform coefficients z = Ax. In fact, for large zero-mean i.i.d sub-Gaussian A, GAMP is characterized by a state evolution whose fixed…

Cited by 0SourceScholar
2015

Matrix Completion from Fewer Entries: Spectral Detectability and Rank Estimation

NeurIPS 2015poster

The completion of low rank matrices from few entries is a task with many practical applications. We consider here two aspects of this problem: detectability, i.e. the ability to estimate the rank $r$ reliably from the fewest possible random entries, and performance in achieving small reconstruction…

Cited by 22SourcePDFScholar