← Search

Leonardo Galli

2 accepted papers

2026

Flatland: The Adventures of Gradient Descent with Large Step Sizes

ICML 2026poster

The training of neural networks often entails objective functions that are not globally $L$-smooth. For these functions, it is both theoretically and practically difficult to reply to the question: what is the largest possible step size that ensures the convergence of gradient descent (GD)? We addre…

Cited by 0SourceScholar
2023

Don't be so Monotone: Relaxing Stochastic Line Search in Over-Parameterized Models

NeurIPS 2023poster

Recent works have shown that line search methods can speed up Stochastic Gradient Descent (SGD) and Adam in modern over-parameterized settings. However, existing line searches may take steps that are smaller than necessary since they require a monotone decrease of the (mini-)batch objective function…

Cited by 10SourcePDFScholar