AAAI 2023technical2 citations

Fast Convergence in Learning Two-Layer Neural Networks with Separable Data

Hossein Taheri, Christos Thrampoulidis

Abstract

Normalized gradient descent has shown substantial success in speeding up the convergence of exponentially-tailed loss functions (which includes exponential and logistic losses) on linear classifiers with separable data. In this paper, we go beyond linear models by studying normalized GD on two-layer neural nets. We prove for exponentially-tailed losses that using normalized GD leads to linear rate of convergence of the training loss to the global optimum. This is made possible by showing certain gradient self-boundedness conditions and a log-Lipschitzness property. We also study generalization of normalized GD for convex objectives via an algorithmic-stability analysis. In particular, we show that normalized GD does not overfit during training by establishing finite-time generalization bounds.

BibTeX
@article{Taheri_Thrampoulidis_2023, title={Fast Convergence in Learning Two-Layer Neural Networks with Separable Data}, volume={37}, url={https://ojs.aaai.org/index.php/AAAI/article/view/26186}, DOI={10.1609/aaai.v37i8.26186}, abstractNote={Normalized gradient descent has shown substantial success in speeding up the convergence of exponentially-tailed loss functions (which includes exponential and logistic losses) on linear classifiers with separable data. In this paper, we go beyond linear models by studying normalized GD on two-layer neural nets. We prove for exponentially-tailed losses that using normalized GD leads to linear rate of convergence of the training loss to the global optimum. This is made possible by showing certain gradient self-boundedness conditions and a log-Lipschitzness property. We also study generalization of normalized GD for convex objectives via an algorithmic-stability analysis. In particular, we show that normalized GD does not overfit during training by establishing finite-time generalization bounds.}, number={8}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Taheri, Hossein and Thrampoulidis, Christos}, year={2023}, month={Jun.}, pages={9944-9952} }
Fast Convergence in Learning Two-Layer Neural Networks with Separable Data · AAAI 2023