2020
Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal Topology
NeurIPS 2020poster
Recent works have shown that gradient descent can find a global minimum for over-parameterized neural networks where the widths of all the hidden layers scale polynomially with N (N being the number of training samples). In this paper, we prove that, for deep networks, a single layer of width N foll…