← Search

Aditya Vardhan Varre

3 accepted papers

2023

On the spectral bias of two-layer linear networks

NeurIPS 2023poster

This paper studies the behaviour of two-layer fully connected networks with linear activations trained with gradient flow on the square loss. We show how the optimization process carries an implicit bias on the parameters that depends on the scale of its initialization. The main result of the paper…

Cited by 15SourcePDFScholar
2023

SGD with Large Step Sizes Learns Sparse Features

ICML 2023poster

We showcase important features of the dynamics of the Stochastic Gradient Descent (SGD) in the training of neural networks. We present empirical observations that commonly used large step sizes (i) may lead the iterates to jump from one side of a valley to the other causing *loss stabilization*, and…

2021

Last iterate convergence of SGD for Least-Squares in the Interpolation regime.

NeurIPS 2021poster

Motivated by the recent successes of neural networks that have the ability to fit the data perfectly \emph{and} generalize well, we study the noiseless model in the fundamental least-squares setup. We assume that an optimum predictor perfectly fits the inputs and outputs $\langle \theta_* , \phi(X)…

Cited by 45SourcePDFScholar