← Search

Alan Milligan

2 accepted papers

2025

Understanding Adam Requires Better Rotation Dependent Assumptions

NeurIPS 2025poster

Despite its widespread adoption, Adam's advantage over Stochastic Gradient Descent (SGD) lacks a comprehensive theoretical explanation. This paper investigates Adam's sensitivity to rotations of the parameter space. We observe that Adam's performance in training transformers degrades under random ro…

Cited by 0SourceScholar
2024

Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

NeurIPS 2024spotlight

Adam has been shown to outperform gradient descent on large language models by a larger margin than on other tasks, but it is unclear why. We show that a key factor in this performance gap is the heavy-tailed class imbalance found in language tasks. When trained with gradient descent, the loss of in…

Cited by 32SourcePDFScholar