← Search

Mitchell A Gordon

1 accepted papers

2021

Data and Parameter Scaling Laws for Neural Machine Translation

EMNLP 2021main

We observe that the development cross-entropy loss of supervised neural machine translation models scales like a power law with the amount of training data and the number of non-embedding parameters in the model. We discuss some practical implications of these results, such as predicting BLEU achiev…