ICLR 2018workshop13 citations
No Spurious Local Minima in a Two Hidden Unit ReLU Network
Chenwei Wu, Jiajun Luo, Jason D. Lee
Abstract
Deep learning models can be efficiently optimized via stochastic gradient descent, but there is little theoretical evidence to support this. A key question in optimization is to understand when the optimization landscape of a neural network is amenable to gradient-based optimization. We focus on a simple neural network two-layer ReLU network with two hidden units, and show that all local minimizers are global. This combined with recent work of Lee et al. (2017); Lee et al. (2016) show that gradient descent converges to the global minimizer.
Non-convex optimizationDeep Learning
BibTeX
@misc{
wu2018no,
title={No Spurious Local Minima in a Two Hidden Unit Re{LU} Network},
author={Chenwei Wu and Jiajun Luo and Jason D. Lee},
year={2018},
url={https://openreview.net/forum?id=B14uJzW0b},
}