← Search

Alex Damian

9 accepted papers

2026

Improved high-dimensional estimation with Langevin dynamics and stochastic weight averaging

ICLR 2026poster

Significant recent work has studied the ability of gradient descent to recover a hidden planted direction $\theta^\star \in S^{d-1}$ in different high-dimensional settings, including tensor PCA and single-index models. The key quantity that governs the ability of gradient descent to traverse these l…

Cited by 0SourceScholar
2025

The Generative Leap: Tight Sample Complexity for Efficiently Learning Gaussian Multi-Index Models

NeurIPS 2025spotlight

In this work we consider generic Gaussian Multi-index models, in which the labels only depend on the (Gaussian) $d$-dimensional inputs through their projection onto a low-dimensional $r = O_d(1)$ subspace, and we study efficient agnostic estimation procedures for this hidden subspace. We introduce t…

Cited by 0SourceScholar
2025

Understanding Optimization in Deep Learning with Central Flows

ICLR 2025poster

Optimization in deep learning remains poorly understood. A key difficulty is that optimizers exhibit complex oscillatory dynamics, referred to as "edge of stability," which cannot be captured by traditional optimization theory. In this paper, we show that the path taken by an oscillatory optimizer…

Cited by 1SourcePDFScholar
2023

Fine-Tuning Language Models with Just Forward Passes

NeurIPS 2023oral

Fine-tuning language models (LMs) has yielded success on diverse downstream tasks, but as LMs grow in size, backpropagation requires a prohibitively large amount of memory. Zeroth-order (ZO) methods can in principle estimate gradients using only two forward passes but are theorized to be catastrophi…

2023

Provable Guarantees for Nonlinear Feature Learning in Three-Layer Neural Networks

NeurIPS 2023spotlight

One of the central questions in the theory of deep learning is to understand how neural networks learn hierarchical features. The ability of deep networks to extract salient features is crucial to both their outstanding generalization ability and the modern deep learning paradigm of pretraining and…

Cited by 21SourcePDFScholar
2023

Self-Stabilization: The Implicit Bias of Gradient Descent at the Edge of Stability

ICLR 2023poster

Traditional analyses of gradient descent show that when the largest eigenvalue of the Hessian, also known as the sharpness $S(\theta)$, is bounded by $2/\eta$, training is "stable" and the training loss decreases monotonically. Recent works, however, have observed that this assumption does not hold…

2023

Smoothing the Landscape Boosts the Signal for SGD: Optimal Sample Complexity for Learning Single Index Models

NeurIPS 2023oral

We focus on the task of learning a single index model $\sigma(w^\star \cdot x)$ with respect to the isotropic Gaussian distribution in $d$ dimensions. Prior work has shown that the sample complexity of learning $w^\star$ is governed by the information exponent $k^\star$ of the link function $\sigma$…

Cited by 51SourcePDFScholar