2019
Leader Stochastic Gradient Descent for Distributed Training of Deep Learning Models
NeurIPS 2019poster
We consider distributed optimization under communication constraints for training deep learning models. We propose a new algorithm, whose parameter updates rely on two forces: a regular gradient step, and a corrective direction dictated by the currently best-performing worker (leader). Our method di…