← Search

Aukosh Jagannath

7 accepted papers

2026

High-dimensional limit theorems for SGD: Momentum and Adaptive Step-sizes

ICLR 2026poster

We develop a high-dimensional scaling limit for Stochastic Gradient Descent with Polyak Momentum (SGD-M) and adaptive step-sizes. This provides a framework to rigourously compare online SGD with some of its popular variants. We show that the scaling limits of SGD-M coincide with those of online SGD…

Cited by 0SourceScholar
2025

Provable Benefits of Unsupervised Pre-training and Transfer Learning via Single-Index Models

ICML 2025spotlight

Unsupervised pre-training and transfer learning are commonly used techniques to initialize training algorithms for neural networks, particularly in settings with limited labeled data. In this paper, we study the effects of unsupervised pre-training and transfer learning on the sample complexity of h…

Cited by 0SourcePDFScholar
2024

High-dimensional SGD aligns with emerging outlier eigenspaces

ICLR 2024spotlight

We rigorously study the joint evolution of training dynamics via stochastic gradient descent (SGD) and the spectra of empirical Hessian and gradient matrices. We prove that in two canonical classification tasks for multi-class high-dimensional mixtures and either 1 or 2-layer neural networks, the SG…

Cited by 11SourcePDFScholar
2023

Optimality of Message-Passing Architectures for Sparse Graphs

NeurIPS 2023poster

We study the node classification problem on feature-decorated graphs in the sparse setting, i.e., when the expected degree of a node is $O(1)$ in the number of nodes, in the fixed-dimensional asymptotic regime, i.e., the dimension of the feature data is fixed while the number of nodes is large. Such…

Cited by 13SourcePDFScholar
2022

High-dimensional limit theorems for SGD: Effective dynamics and critical scaling

NeurIPS 2022accept

We study the scaling limits of stochastic gradient descent (SGD) with constant step-size in the high-dimensional regime. We prove limit theorems for the trajectories of summary statistics (i.e., finite-dimensional functions) of SGD as the dimension goes to infinity. Our approach allows one to choose…

Cited by 87SourcePDFScholar
2021

Graph Convolution for Semi-Supervised Classification: Improved Linear Separability and Out-of-Distribution Generalization

ICML 2021spotlight

Recently there has been increased interest in semi-supervised classification in the presence of graphical information. A new class of learning models has emerged that relies, at its most basic level, on classifying the data after first applying a graph convolution. To understand the merits of this a…

Cited by 93SourcePDFScholar