← Search

Scott Pesme

9 accepted papers

2025

A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation

NeurIPS 2025poster

We study the dynamics of gradient flow with small weight decay on general training losses $F: \mathbb{R}^d \to \mathbb{R}$. Under mild regularity assumptions and assuming convergence of the unregularised gradient flow, we show that the trajectory with weight decay $\lambda$ exhibits a two-phase beha…

Cited by 0SourceScholar
2025

MAP Estimation with Denoisers: Convergence Rates and Guarantees

NeurIPS 2025poster

Denoiser models have become powerful tools for inverse problems, enabling the use of pretrained networks to approximate the score of a smoothed prior distribution. These models are often used in heuristic iterative schemes aimed at solving Maximum a Posteriori (MAP) optimisation problems, where the…

Cited by 0SourceScholar
2024

Leveraging Continuous Time to Understand Momentum When Training Diagonal Linear Networks

AISTATS 2024poster

In this work, we investigate the effect of momentum on the optimisation trajectory of gradient descent. We leverage a continuous-time approach in the analysis of momentum gradient descent with step size $\gamma$ and momentum parameter $\beta$ that allows us to identify an intrinsic quantity $\lambda…

Cited by 12SourcePDFScholar
2023

(S)GD over Diagonal Linear Networks: Implicit bias, Large Stepsizes and Edge of Stability

NeurIPS 2023poster

In this paper, we investigate the impact of stochasticity and large stepsizes on the implicit regularisation of gradient descent (GD) and stochastic gradient descent (SGD) over $2$-layer diagonal linear networks. We prove the convergence of GD and SGD with macroscopic stepsizes in an overparametrise…

Cited by 35SourcePDFScholar
2021

Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of Stochasticity

NeurIPS 2021poster

Understanding the implicit bias of training algorithms is of crucial importance in order to explain the success of overparametrised neural networks. In this paper, we study the dynamics of stochastic gradient descent over diagonal linear networks through its continuous time version, namely stochasti…

Cited by 127SourcePDFScholar
2020

On Convergence-Diagnostic based Step Sizes for Stochastic Gradient Descent

ICML 2020poster

Constant step-size Stochastic Gradient Descent exhibits two phases: a transient phase during which iterates make fast progress towards the optimum, followed by a stationary phase during which iterates oscillate around the optimal point. In this paper, we show that efficiently detecting this transiti…

Cited by 27SourcePDFScholar