2022
S2 Reducer: High-Performance Sparse Communication to Accelerate Distributed Deep Learning
ICASSP 2022accepted
Distributed stochastic gradient descent (SGD) approach has been widely used in large-scale deep learning, and the gradient collective method is vital to ensure the training scalability of the distributed deep learning system. Collective communication such as AllReduce has been widely adopted for the…