2024
Data, Data Everywhere: A Guide for Pretraining Dataset Construction
EMNLP 2024main
The impressive capabilities of recent language models can be largely attributed to the multi-trillion token pretraining datasets that they are trained on. However, model developers fail to disclose their construction methodology which has lead to a lack of open information on how to develop effectiv…