← Search

Kaiyu Shi

1 accepted papers

2021

Memory-Efficient Pipeline-Parallel DNN Training

ICML 2021spotlight

Many state-of-the-art ML results have been obtained by scaling up the number of parameters in existing models. However, parameters and activations for such large models often do not fit in the memory of a single accelerator device; this means that it is necessary to distribute training of large mode…

Cited by 274SourcePDFScholar