2026
Flatland: The Adventures of Gradient Descent with Large Step Sizes
ICML 2026poster
The training of neural networks often entails objective functions that are not globally $L$-smooth. For these functions, it is both theoretically and practically difficult to reply to the question: what is the largest possible step size that ensures the convergence of gradient descent (GD)? We addre…