Geometric Convergence of Gauss–Newton for Neural Networks: Riemannian Geometry and Adaptive Damping
Ill-conditioned kernel matrices and loss landscapes can make first-order methods for training neural networks converge slowly. We establish non-asymptotic convergence bounds for the Gauss–Newton method in both under- and overparameterized regimes, showing it avoids these conditioning bottlenecks. In…