← Search

Alireza Mousavi-Hosseini

10 accepted papers

2025

From Information to Generative Exponent: Learning Rate Induces Phase Transitions in SGD

NeurIPS 2025poster

To understand feature learning dynamics in neural networks, recent theoretical works have focused on gradient-based learning of Gaussian single-index models, where the label is a nonlinear function of a latent one-dimensional projection of the input. While the sample complexity of online SGD is dete…

Cited by 0SourceScholar
2025

Learning Multi-Index Models with Neural Networks via Mean-Field Langevin Dynamics

ICLR 2025poster

We study the problem of learning multi-index models in high-dimensions using a two-layer neural network trained with the mean-field Langevin algorithm. Under mild distributional assumptions on the data, we characterize the effective dimension $d_{\mathrm{eff}}$ that controls both sample and computat…

Cited by 3SourcePDFScholar
2025

Robust Feature Learning for Multi-Index Models in High Dimensions

ICLR 2025poster

Recently, there have been numerous studies on feature learning with neural networks, specifically on learning single- and multi-index models where the target is a function of a low-dimensional projection of the input. Prior works have shown that in high dimensions, the majority of the compute and da…

2025

When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical Perspective

NeurIPS 2025poster

Theoretical efforts to prove advantages of Transformers in comparison with classical architectures such as feedforward and recurrent neural networks have mostly focused on representational power. In this work, we take an alternative perspective and prove that even with infinite compute, feedforward…

Cited by 0SourcecodeScholar
2024

A Separation in Heavy-Tailed Sampling: Gaussian vs. Stable Oracles for Proximal Samplers

NeurIPS 2024poster

We study the complexity of heavy-tailed sampling and present a separation result in terms of obtaining high-accuracy versus low-accuracy guarantees i.e., samplers that require only $\mathcal{O}(\log(1/\varepsilon))$ versus $\Omega(\text{poly}(1/\varepsilon))$ iterations to output a sample which is $…

Cited by 2SourcePDFScholar
2024

Mean-Field Langevin Dynamics for Signed Measures via a Bilevel Approach

NeurIPS 2024spotlight

Mean-field Langevin dynamics (MLFD) is a class of interacting particle methods that tackle convex optimization over probability measures on a manifold, which are scalable, versatile, and enjoy computational guarantees. However, some important problems -- such as risk minimization for infinite width…

2023

Gradient-Based Feature Learning under Structured Data

NeurIPS 2023poster

Recent works have demonstrated that the sample complexity of gradient-based learning of single index models, i.e. functions that depend on a 1-dimensional projection of the input data, is governed by their information exponent. However, these results are only concerned with isotropic data, while in…

Cited by 29SourcePDFScholar
2023

Neural Networks Efficiently Learn Low-Dimensional Representations with SGD

ICLR 2023top-25%

We study the problem of training a two-layer neural network (NN) of arbitrary width using stochastic gradient descent (SGD) where the input $\boldsymbol{x}\in \mathbb{R}^d$ is Gaussian and the target $y \in \mathbb{R}$ follows a multiple-index model, i.e., $y=g(\langle\boldsymbol{u_1},\boldsymbol{x}…

Cited by 71SourcePDFScholar