2025
A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation
NeurIPS 2025poster
We study the dynamics of gradient flow with small weight decay on general training losses $F: \mathbb{R}^d \to \mathbb{R}$. Under mild regularity assumptions and assuming convergence of the unregularised gradient flow, we show that the trajectory with weight decay $\lambda$ exhibits a two-phase beha…