2020
On the Generalization Benefit of Noise in Stochastic Gradient Descent
ICML 2020poster
It has long been argued that minibatch stochastic gradient descent can generalize better than large batch gradient descent in deep neural networks. However recent papers have questioned this claim, arguing that this effect is simply a consequence of suboptimal hyperparameter tuning or insufficient c…