2020
Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks
IJCAI 2020poster
Adaptive gradient methods, which adopt historical gradient information to automatically adjust the learning rate, despite the nice property of fast convergence, have been observed to generalize worse than stochastic gradient descent (SGD) with momentum in training deep neural networks. This leaves h…