2024
MiLe Loss: a New Loss for Mitigating the Bias of Learning Difficulties in Generative Language Models
NAACL 2024findings
Generative language models are usually pre-trained on large text corpus via predicting the next token (i.e., sub-word/word/phrase) given the previous ones. Recent works have demonstrated the impressive performance of large generative language models on downstream tasks. However, existing generative…