2023
Birder: Communication-Efficient 1-bit Adaptive Optimizer for Practical Distributed DNN Training
NeurIPS 2023poster
Various gradient compression algorithms have been proposed to alleviate the communication bottleneck in distributed learning, and they have demonstrated effectiveness in terms of high compression ratios and theoretical low communication complexity. However, when it comes to practically training mod…