NeurIPS 2018poster24 citations

The Effect of Network Width on the Performance of Large-batch Training

Lingjiao Chen, Hongyi Wang, Jinman Zhao, Dimitris Papailiopoulos, Paraschos Koutris

Abstract

Distributed implementations of mini-batch stochastic gradient descent (SGD) suffer from communication overheads, attributed to the high frequency of gradient updates inherent in small-batch training. Training with large batches can reduce these overheads; however it besets the convergence of the algorithm and the generalization performance.

BibTeX
@inproceedings{NEURIPS2018_e7c573c1,
 author = {Chen, Lingjiao and Wang, Hongyi and Zhao, Jinman and Papailiopoulos, Dimitris and Koutris, Paraschos},
 booktitle = {Advances in Neural Information Processing Systems},
 editor = {S. Bengio and H. Wallach and H. Larochelle and K. Grauman and N. Cesa-Bianchi and R. Garnett},
 pages = {},
 publisher = {Curran Associates, Inc.},
 title = {The Effect of Network Width on the Performance of  Large-batch Training},
 url = {https://proceedings.neurips.cc/paper_files/paper/2018/file/e7c573c14a09b84f6b7782ce3965f335-Paper.pdf},
 volume = {31},
 year = {2018}
}
The Effect of Network Width on the Performance of Large-batch Training · NeurIPS 2018