2022
Eliminating Sharp Minima from SGD with Truncated Heavy-tailed Noise
ICLR 2022poster
The empirical success of deep learning is often attributed to SGD’s mysterious ability to avoid sharp local minima in the loss landscape, as sharp minima are known to lead to poor generalization. Recently, empirical evidence of heavy-tailed gradient noise was reported in many deep learning tasks; a…