2022
DRAGONN: Distributed Randomized Approximate Gradients of Neural Networks
ICML 2022spotlight
Data-parallel distributed training (DDT) has become the de-facto standard for accelerating the training of most deep learning tasks on massively parallel hardware. In the DDT paradigm, the communication overhead of gradient synchronization is the major efficiency bottleneck. A widely adopted approac…