2020
Towards Better Generalization of Adaptive Gradient Methods
NeurIPS 2020poster
Adaptive gradient methods such as AdaGrad, RMSprop and Adam have been optimizers of choice for deep learning due to their fast training speed. However, it was recently observed that their generalization performance is often worse than that of SGD for over-parameterized neural networks. While new alg…