Global Convergence of Gradient Descent for Deep Linear Residual Networks
Abstract
We analyze the global convergence of gradient descent for deep linear residual networks by proposing a new initialization: zero-asymmetric (ZAS) initialization. It is motivated by avoiding stable manifolds of saddle points. We prove that under the ZAS initialization, for an arbitrary target matrix, gradient descent converges to an $\varepsilon$-optimal point in $O\left( L^3 \log(1/\varepsilon) \right)$ iterations, which scales polynomially with the network depth $L$. Our result and the $\exp(\Omega(L))$ convergence time for the standard initialization (Xavier or near-identity) \cite{shamir2018exponential} together demonstrate the importance of the residual structure and the initialization in the optimization for deep linear neural networks, especially when $L$ is large.
BibTeX
@inproceedings{NEURIPS2019_14da15db,
author = {Wu, Lei and Wang, Qingcan and Ma, Chao},
booktitle = {Advances in Neural Information Processing Systems},
editor = {H. Wallach and H. Larochelle and A. Beygelzimer and F. d\textquotesingle Alch\'{e}-Buc and E. Fox and R. Garnett},
pages = {},
publisher = {Curran Associates, Inc.},
title = {Global Convergence of Gradient Descent for Deep Linear Residual Networks},
url = {https://proceedings.neurips.cc/paper_files/paper/2019/file/14da15db887a4b50efe5c1bc66537089-Paper.pdf},
volume = {32},
year = {2019}
}