← Search

Mohamad Amin Mohamadi

4 accepted papers

2025

Adam Exploits $\ell_\infty$-geometry of Loss Landscape via Coordinate-wise Adaptivity

ICLR 2025spotlight

Adam outperforms SGD when training language models. Yet this advantage is not well-understood theoretically -- previous convergence analysis for Adam and SGD mainly focuses on the number of steps $T$ and is already minimax-optimal in non-convex cases, which are both $\widetilde{O}(T^{-1/4})$. In th…

2024

Why Do You Grok? A Theoretical Analysis on Grokking Modular Addition

ICML 2024poster

We present a theoretical explanation of the “grokking” phenomenon (Power et al., 2022), where a model generalizes long after overfitting, for the originally-studied problem of modular addition. First, we show that early in gradient descent, so that the “kernel regime” approximately holds, no permuta…

Cited by 7SourcePDFScholar
2023

A Fast, Well-Founded Approximation to the Empirical Neural Tangent Kernel

ICML 2023poster

Empirical neural tangent kernels (eNTKs) can provide a good understanding of a given network's representation: they are often far less expensive to compute and applicable more broadly than infinite-width NTKs. For networks with $O$ output units (e.g. an $O$-class classifier), however, the eNTK on $N…

Cited by 27SourcePDFScholar
2022

Making Look-Ahead Active Learning Strategies Feasible with Neural Tangent Kernels

NeurIPS 2022accept

We propose a new method for approximating active learning acquisition strategies that are based on retraining with hypothetically-labeled candidate data points. Although this is usually infeasible with deep networks, we use the neural tangent kernel to approximate the result of retraining, and prove…

Cited by 32SourcePDFScholar