2020
Bounds on Over-Parameterization for Guaranteed Existence of Descent Paths in Shallow ReLU Networks
ICLR 2020poster
We study the landscape of squared loss in neural networks with one-hidden layer and ReLU activation functions. Let $m$ and $d$ be the widths of hidden and input layers, respectively. We show that there exist poor local minima with positive curvature for some training sets of size $n\geq m+2d-2$. By…