NeurIPS 2021poster4 citations
Faster Directional Convergence of Linear Neural Networks under Spherically Symmetric Data
Dachao Lin, Ruoyu Sun, Zhihua Zhang
Abstract
In this paper, we study gradient methods for training deep linear neural networks with binary cross-entropy loss. In particular, we show global directional convergence guarantees from a polynomial rate to a linear rate for (deep) linear networks with spherically symmetric data distribution, which can be viewed as a specific zero-margin dataset. Our results do not require the assumptions in other works such as small initial loss, presumed convergence of weight direction, or overparameterization. We also characterize our findings in experiments.
optimization for deep linear networksglobal directional convergence
BibTeX
@inproceedings{
lin2021faster,
title={Faster Directional Convergence of Linear Neural Networks under Spherically Symmetric Data},
author={Dachao Lin and Ruoyu Sun and Zhihua Zhang},
booktitle={Advances in Neural Information Processing Systems},
editor={A. Beygelzimer and Y. Dauphin and P. Liang and J. Wortman Vaughan},
year={2021},
url={https://openreview.net/forum?id=Q9hZdUBTC9S}
}