2023
Coordinating Distributed Example Orders for Provably Accelerated Training
NeurIPS 2023poster
Recent research on online Gradient Balancing (GraB) has revealed that there exist permutation-based example orderings for SGD that are guaranteed to outperform random reshuffling (RR). Whereas RR arbitrarily permutes training examples, GraB leverages stale gradients from prior epochs to order exampl…