Convergence Rates for Gradient Descent on the Edge of Stability for Overparametrised Least Squares
Classical optimisation theory guarantees monotonic objective decrease for gradient descent (GD) when employed in a small step size, or "stable", regime. In contrast, gradient descent on neural networks is frequently performed in a large step size regime called the "edge of stability", in which the o…