← Search

Dimitris Achlioptas

1 accepted papers

2020

Bad Global Minima Exist and SGD Can Reach Them

NeurIPS 2020poster

Several works have aimed to explain why overparameterized neural networks generalize well when trained by Stochastic Gradient Descent (SGD). The consensus explanation that has emerged credits the randomized nature of SGD for the bias of the training process towards low-complexity models and, thus,…