2025
Escaping saddle points without Lipschitz smoothness: the power of nonlinear preconditioning
NeurIPS 2025spotlight
We study generalized smoothness in nonconvex optimization, focusing on $(L_0, L_1)$-smoothness and anisotropic smoothness. The former was empirically derived from practical neural network training examples, while the latter arises naturally in the analysis of nonlinearly preconditioned gradient meth…