← Search

Courtney Paquette

9 accepted papers

2025

Dimension-adapted Momentum Outscales SGD

NeurIPS 2025spotlight

We investigate scaling laws for stochastic momentum algorithms on the power law random features model, parameterized by data complexity, target complexity, and model size. When trained with a stochastic momentum algorithm, our analysis reveals four distinct loss curve shapes determined by varying da…

Cited by 0SourceScholar
2025

Implicit Diffusion: Efficient optimization through stochastic sampling

AISTATS 2025oral

Sampling and automatic differentiation are both ubiquitous in modern machine learning. At its intersection, differentiating through a sampling operation, with respect to the parameters of the sampling process, is a problem that is both challenging and broadly applicable. We introduce a general frame…

Cited by 0SourceScholar
2024

4+3 Phases of Compute-Optimal Neural Scaling Laws

NeurIPS 2024spotlight

We consider the solvable neural scaling model with three parameters: data complexity, target complexity, and model-parameter-count. We use this neural scaling model to derive new predictions about the compute-limited, infinite-data scaling law regime. To train the neural scaling model, we run one-p…

Cited by 18SourcePDFScholar
2024

The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate Algorithms

NeurIPS 2024poster

We develop a framework for analyzing the training and learning rate dynamics on a large class of high-dimensional optimization problems, which we call the high line, trained using one-pass stochastic gradient descent (SGD) with adaptive learning rates. We give exact expressions for the risk and lear…

2022

Implicit Regularization or Implicit Conditioning? Exact Risk Trajectories of SGD in High Dimensions

NeurIPS 2022accept

Stochastic gradient descent (SGD) is a pillar of modern machine learning, serving as the go-to optimization algorithm for a diverse array of problems. While the empirical success of SGD is often attributed to its computational efficiency and favorable generalization behavior, neither effect is well…

Cited by 19SourcePDFScholar
2022

Only tails matter: Average-Case Universality and Robustness in the Convex Regime

ICML 2022spotlight

The recently developed average-case analysis of optimization methods allows a more fine-grained and representative convergence analysis than usual worst-case results. In exchange, this analysis requires a more precise hypothesis over the data generating process, namely assuming knowledge of the expe…

Cited by 11SourcePDFScholar
2022

Trajectory of Mini-Batch Momentum: Batch Size Saturation and Convergence in High Dimensions

NeurIPS 2022accept

We analyze the dynamics of large batch stochastic gradient descent with momentum (SGD+M) on the least squares problem when both the number of samples and dimensions are large. In this setting, we show that the dynamics of SGD+M converge to a deterministic discrete Volterra equation as dimension incr…

Cited by 20SourcePDFScholar
2018

Catalyst for Gradient-based Nonconvex Optimization

AISTATS 2018poster

We introduce a generic scheme to solve nonconvex optimization problems using gradient-based algorithms originally designed for minimizing convex functions. Even though these methods may originally require convexity to operate, the proposed approach allows one to use them without assuming any knowled…

Cited by 0SourcePDFScholar