2018
A Linear Speedup Analysis of Distributed Deep Learning with Sparse and Quantized Communication
NeurIPS 2018poster
The large communication overhead has imposed a bottleneck on the performance of distributed Stochastic Gradient Descent (SGD) for training deep neural networks. Previous works have demonstrated the potential of using gradient sparsification and quantization to reduce the communication cost. Howeve…