← Search

Salma Tarmoun

5 accepted papers

2025

Convergence Rates for Gradient Descent on the Edge of Stability for Overparametrised Least Squares

NeurIPS 2025poster

Classical optimisation theory guarantees monotonic objective decrease for gradient descent (GD) when employed in a small step size, or "stable", regime. In contrast, gradient descent on neural networks is frequently performed in a large step size regime called the "edge of stability", in which the o…

Cited by 0SourceScholar
2025

Understanding the Learning Dynamics of LoRA: A Gradient Flow Perspective on Low-Rank Adaptation in Matrix Factorization

AISTATS 2025poster

Despite the empirical success of Low-Rank Adaptation (LoRA) in fine-tuning pre-trained models, there is little theoretical understanding of how first-order methods with carefully crafted initialization adapt models to new tasks. In this work, we take the first step towards bridging this gap by theor…

Cited by 0SourceScholar
2023

Linear Convergence of Gradient Descent For Finite Width Over-parametrized Linear Networks With General Initialization

AISTATS 2023poster

Recent theoretical analyses of the convergence of gradient descent (GD) to a global minimum for over-parametrized neural networks make strong assumptions on the step size (infinitesimal), the hidden-layer width (infinite), or the initialization (spectral, balanced). In this work, we relax these assu…

Cited by 8SourcePDFScholar
2021

On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear Networks

ICML 2021spotlight

Neural networks trained via gradient descent with random initialization and without any regularization enjoy good generalization performance in practice despite being highly overparametrized. A promising direction to explain this phenomenon is to study how initialization and overparametrization affe…

Cited by 61SourcePDFScholar
2021

Understanding the Dynamics of Gradient Flow in Overparameterized Linear models

ICML 2021spotlight

We provide a detailed analysis of the dynamics ofthe gradient flow in overparameterized two-layerlinear models. A particularly interesting featureof this model is that its nonlinear dynamics can beexactly solved as a consequence of a large num-ber of conservation laws that constrain the systemto fol…

Cited by 71SourcePDFScholar