2021
Moshpit SGD: Communication-Efficient Decentralized Training on Heterogeneous Unreliable Devices
NeurIPS 2021poster
Training deep neural networks on large datasets can often be accelerated by using multiple compute nodes. This approach, known as distributed training, can utilize hundreds of computers via specialized message-passing protocols such as Ring All-Reduce. However, running these protocols at scale requ…