← Search

Jinxing Yu

1 accepted papers

2020

Towards Better Generalization of Adaptive Gradient Methods

NeurIPS 2020poster

Adaptive gradient methods such as AdaGrad, RMSprop and Adam have been optimizers of choice for deep learning due to their fast training speed. However, it was recently observed that their generalization performance is often worse than that of SGD for over-parameterized neural networks. While new alg…

Cited by 25SourcePDFScholar