NeurIPS 2020poster36 citations
On Convergence and Generalization of Dropout Training
Abstract
We study dropout in two-layer neural networks with rectified linear unit (ReLU) activations. Under mild overparametrization and assuming that the limiting kernel can separate the data distribution with a positive margin, we show that the dropout training with logistic loss achieves $\epsilon$-suboptimality in the test error in $O(1/\epsilon)$ iterations.
BibTeX
@inproceedings{NEURIPS2020_f1de5100,
author = {Mianjy, Poorya and Arora, Raman},
booktitle = {Advances in Neural Information Processing Systems},
editor = {H. Larochelle and M. Ranzato and R. Hadsell and M.F. Balcan and H. Lin},
pages = {21151--21161},
publisher = {Curran Associates, Inc.},
title = {On Convergence and Generalization of Dropout Training},
url = {https://proceedings.neurips.cc/paper_files/paper/2020/file/f1de5100906f31712aaa5166689bfdf4-Paper.pdf},
volume = {33},
year = {2020}
}