LegoMT2: Selective Asynchronous Sharded Data Parallel Training for Massive Neural Machine Translation
It is a critical challenge to learn a single model for massive languages. Prior methods focus on increasing the model size and training data size. However, large models are difficult to optimize efficiently even with distributed parallel training and translation capacity can interfere among language…