← Search

Radu-Alexandru Dragomir

2 accepted papers

2025

A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation

NeurIPS 2025poster

We study the dynamics of gradient flow with small weight decay on general training losses $F: \mathbb{R}^d \to \mathbb{R}$. Under mild regularity assumptions and assuming convergence of the unregularised gradient flow, we show that the trajectory with weight decay $\lambda$ exhibits a two-phase beha…

Cited by 0SourceScholar