2021
Data and Parameter Scaling Laws for Neural Machine Translation
EMNLP 2021main
We observe that the development cross-entropy loss of supervised neural machine translation models scales like a power law with the amount of training data and the number of non-embedding parameters in the model. We discuss some practical implications of these results, such as predicting BLEU achiev…