← Search

Hugo Cui

14 accepted papers

2026

High-Dimensional Analysis of Single-Layer Attention for Sparse-Token Classification

ICLR 2026poster

When and how can an attention mechanism learn to selectively attend to informative tokens, thereby enabling detection of weak, rare, and sparsely located features? We address these questions theoretically in a sparse-token classification model in which positive samples embed a weak signal vector in…

Cited by 0SourceScholar
2025

A Random Matrix Theory Perspective on the Spectrum of Learned Features and Asymptotic Generalization Capabilities

AISTATS 2025oral

A key property of neural networks is their capacity of adapting to data during training. Yet, our current mathematical understanding of feature learning and its relationship to generalization remain limited. In this work, we provide a random matrix analysis of how fully-connected two-layer neural ne…

Cited by 0SourceScholar
2025

Fundamental limits of learning in sequence multi-index models and deep attention networks: high-dimensional asymptotics and sharp thresholds

ICML 2025poster

In this manuscript, we study the learning of deep attention neural networks, defined as the composition of multiple self-attention layers, with tied and low-rank weights. We first establish a mapping of such models to sequence multi-index models, a generalization of the widely studied multi-index m…

2024

A Phase Transition between Positional and Semantic Learning in a Solvable Model of Dot-Product Attention

NeurIPS 2024spotlight

Many empirical studies have provided evidence for the emergence of algorithmic mechanisms (abilities) in the learning of language models, that lead to qualitative improvements of the model capabilities. Yet, a theoretical characterization of how such mechanisms emerge remains elusive. In this paper,…

Cited by 13SourcePDFScholar
2024

Analysis of Learning a Flow-based Generative Model from Limited Sample Complexity

ICLR 2024poster

We study the problem of training a flow-based generative model, parametrized by a two-layer autoencoder, to sample from a high-dimensional Gaussian mixture. We provide a sharp end-to-end analysis of the problem. First, we provide a tight closed-form characterization of the learnt velocity field, whe…

2024

Asymptotics of Learning with Deep Structured (Random) Features

ICML 2024poster

For a large class of feature maps we provide a tight asymptotic characterisation of the test error associated with learning the readout layer, in the high-dimensional limit where the input dimension, hidden layer widths, and number of training samples are proportionally large. This characterization…

2024

Asymptotics of feature learning in two-layer networks after one gradient-step

ICML 2024spotlight

In this manuscript, we investigate the problem of how two-layer neural networks learn features from data, and improve over the kernel regime, after being trained with a single gradient descent step. Leveraging the insight from (Ba et al., 2022), we model the trained network by a spiked Random Featur…

2023

Deterministic equivalent and error universality of deep random features learning

ICML 2023poster

This manuscript considers the problem of learning a random Gaussian network function using a fully connected network with frozen intermediate layers and trainable readout layer. This problem can be seen as a natural generalization of the widely studied random features model to deeper architectures.…

2021

Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy Regime

NeurIPS 2021poster

In this manuscript we consider Kernel Ridge Regression (KRR) under the Gaussian design. Exponents for the decay of the excess generalization error of KRR have been reported in various works under the assumption of power-law decay of eigenvalues of the features co-variance. These decays were, however…

Cited by 108SourcePDFScholar
2021

Learning curves of generic features maps for realistic datasets with a teacher-student model

NeurIPS 2021poster

Teacher-student models provide a framework in which the typical-case performance of high-dimensional supervised learning can be described in closed form. The assumptions of Gaussian i.i.d. input data underlying the canonical teacher-student model may, however, be perceived as too restrictive to capt…