← Search

Yatin Dandi

12 accepted papers

2026

Efficient Learning of Compositional Targets with Hierarchical Spectral Methods

ICML 2026poster

Why depth yields a genuine computational advantage over shallow methods remains a central open question in learning theory. We study this question in a controlled high-dimensional Gaussian setting, focusing on compositional target functions. We analyze their learnability using an explicit three-laye…

Cited by 0SourceScholar
2025

A Random Matrix Theory Perspective on the Spectrum of Learned Features and Asymptotic Generalization Capabilities

AISTATS 2025oral

A key property of neural networks is their capacity of adapting to data during training. Yet, our current mathematical understanding of feature learning and its relationship to generalization remain limited. In this work, we provide a random matrix analysis of how fully-connected two-layer neural ne…

Cited by 0SourceScholar
2025

Fundamental computational limits of weak learnability in high-dimensional multi-index models

AISTATS 2025poster

Multi-index models - functions which only depend on the covariates through a non-linear transformation of their projection on a subspace - are a useful benchmark for investigating feature learning with neural networks. This paper examines the theoretical boundaries of efficient learnability in this…

Cited by 0SourcecodeScholar
2025

Fundamental limits of learning in sequence multi-index models and deep attention networks: high-dimensional asymptotics and sharp thresholds

ICML 2025poster

In this manuscript, we study the learning of deep attention neural networks, defined as the composition of multiple self-attention layers, with tied and low-rank weights. We first establish a mapping of such models to sequence multi-index models, a generalization of the widely studied multi-index m…

2025

Optimal Spectral Transitions in High-Dimensional Multi-Index Models

NeurIPS 2025poster

We consider the problem of how many samples from a Gaussian multi-index model are required to weakly reconstruct the relevant index subspace. Despite its increasing popularity as a testbed for investigating the computational complexity of neural networks, results beyond the single-index setting rema…

Cited by 0SourceScholar
2025

The Computational Advantage of Depth in Learning High-Dimensional Hierarchical Targets

NeurIPS 2025spotlight

Understanding the advantages of deep neural networks trained by gradient descent (GD) compared to shallow models remains an open theoretical challenge. In this paper, we introduce a class of target functions (single and multi-index Gaussian hierarchical targets) that incorporate a hierarchy of laten…

Cited by 0SourceScholar
2024

Asymptotics of feature learning in two-layer networks after one gradient-step

ICML 2024spotlight

In this manuscript, we investigate the problem of how two-layer neural networks learn features from data, and improve over the kernel regime, after being trained with a single gradient descent step. Leveraging the insight from (Ba et al., 2022), we model the trained network by a spiked Random Featur…

2024

Online Learning and Information Exponents: The Importance of Batch size & Time/Complexity Tradeoffs

ICML 2024poster

We study the impact of the batch size $n_b$ on the iteration time $T$ of training two-layer neural networks with one-pass stochastic gradient descent (SGD) on multi-index target functions of isotropic covariates. We characterize the optimal batch size minimizing the iteration time as a function of t…

Cited by 5SourcePDFScholar
2024

The Benefits of Reusing Batches for Gradient Descent in Two-Layer Networks: Breaking the Curse of Information and Leap Exponents

ICML 2024poster

We investigate the training dynamics of two-layer neural networks when learning multi-index target functions. We focus on multi-pass gradient descent (GD) that reuses the batches multiple times and show that it significantly changes the conclusion about which functions are learnable compared to sing…

2023

Universality laws for Gaussian mixtures in generalized linear models

NeurIPS 2023poster

A recent line of work in high-dimensional statistics working under the Gaussian mixture hypothesis has led to a number of results in the context of empirical risk minimization, Bayesian uncertainty quantification, separation of kernel methods and neural networks, ensembling and fluctuation of random…

Cited by 29SourcePDFScholar