← Search

Elliot Paquette

8 accepted papers

2025

Dimension-adapted Momentum Outscales SGD

NeurIPS 2025spotlight

We investigate scaling laws for stochastic momentum algorithms on the power law random features model, parameterized by data complexity, target complexity, and model size. When trained with a stochastic momentum algorithm, our analysis reveals four distinct loss curve shapes determined by varying da…

Cited by 0SourceScholar
2025

Exact risk curves of signSGD in High-Dimensions: quantifying preconditioning and noise-compression effects

ICML 2025poster

In recent years, SignSGD has garnered interest as both a practical optimizer as well as a simple model to understand adaptive optimizers like Adam. Though there is a general consensus that SignSGD acts to precondition optimization and reshapes noise, quantitatively understanding these effects in the…

Cited by 1SourcePDFScholar
2025

To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions

ICLR 2025poster

The success of modern machine learning is due in part to the adaptive optimization methods that have been developed to deal with the difficulties of training large models over complex datasets. One such method is gradient clipping: a practical procedure with limited theoretical underpinnings. In thi…

2024

4+3 Phases of Compute-Optimal Neural Scaling Laws

NeurIPS 2024spotlight

We consider the solvable neural scaling model with three parameters: data complexity, target complexity, and model-parameter-count. We use this neural scaling model to derive new predictions about the compute-limited, infinite-data scaling law regime. To train the neural scaling model, we run one-p…

Cited by 18SourcePDFScholar
2024

The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate Algorithms

NeurIPS 2024poster

We develop a framework for analyzing the training and learning rate dynamics on a large class of high-dimensional optimization problems, which we call the high line, trained using one-pass stochastic gradient descent (SGD) with adaptive learning rates. We give exact expressions for the risk and lear…

2022

Implicit Regularization or Implicit Conditioning? Exact Risk Trajectories of SGD in High Dimensions

NeurIPS 2022accept

Stochastic gradient descent (SGD) is a pillar of modern machine learning, serving as the go-to optimization algorithm for a diverse array of problems. While the empirical success of SGD is often attributed to its computational efficiency and favorable generalization behavior, neither effect is well…

Cited by 19SourcePDFScholar
2022

Trajectory of Mini-Batch Momentum: Batch Size Saturation and Convergence in High Dimensions

NeurIPS 2022accept

We analyze the dynamics of large batch stochastic gradient descent with momentum (SGD+M) on the least squares problem when both the number of samples and dimensions are large. In this setting, we show that the dynamics of SGD+M converge to a deterministic discrete Volterra equation as dimension incr…

Cited by 20SourcePDFScholar