← Search

Behrad Moniri

6 accepted papers

2025

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning

ICML 2025poster

Layer-wise preconditioning methods are a family of memory-efficient optimization algorithms that introduce preconditioners per axis of each layer's weight tensors. These methods have seen a recent resurgence, demonstrating impressive performance relative to entry-wise ("diagonal") preconditioning me…

Cited by 0SourcePDFScholar
2024

A Theory of Non-Linear Feature Learning with One Gradient Step in Two-Layer Neural Networks

ICML 2024poster

Feature learning is thought to be one of the fundamental reasons for the success of deep neural networks. It is rigorously known that in two-layer fully-connected neural networks under certain conditions, one step of gradient descent on the first layer can lead to feature learning; characterized by…

Cited by 31SourcePDFScholar
2023

Demystifying Disagreement-on-the-Line in High Dimensions

ICML 2023poster

Evaluating the performance of machine learning models under distribution shifts is challenging, especially when we only have unlabeled data from the shifted (target) domain, along with labeled data from the original (source) domain. Recent work suggests that the notion of *disagreement*, the degree…

2021

Rate-Distortion Analysis of Minimum Excess Risk in Bayesian Learning

ICML 2021oral

In parametric Bayesian learning, a prior is assumed on the parameter $W$ which determines the distribution of samples. In this setting, Minimum Excess Risk (MER) is defined as the difference between the minimum expected loss achievable when learning from data and the minimum expected loss that could…

Cited by 12SourcePDFScholar