← Search

Mikhail Smelyanskiy

1 accepted papers

2017

On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima

ICLR 2017oral

The stochastic gradient descent (SGD) method and its variants are algorithms of choice for many Deep Learning tasks. These methods operate in a small-batch regime wherein a fraction of the training data, say $32$--$512$ data points, is sampled to compute an approximation to the gradient. It has bee…

Cited by 4006SourcecodeScholar