Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu
Abstract
Information about pretraining corpora used to train the current best-performing language models is seldom discussed: commercial models rarely detail their data, and even open models are often released without accompanying training data or recipes to reproduce them. As a result, it is challenging to conduct and advance scientific research on language modeling, such as understanding how training data impacts model capabilities and limitations. To facilitate scientific research on language model pretraining, we curate and release Dolma, a three-trillion-token English corpus, built from a diverse mixture of web content, scientific papers, code, public-domain books, social media, and encyclopedic materials. We extensively document Dolma, including its design principles, details about its construction, and a summary of its contents. We present analyses and experimental results on intermediate states of Dolma to share what we have learned about important data curation practices. Finally, we open-source our data curation toolkit to enable reproduction of our work as well as support further research in large-scale data curation.
BibTeX
@inproceedings{soldaini-etal-2024-dolma,
title = "Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research",
author = "Soldaini, Luca and
Kinney, Rodney and
Bhagia, Akshita and
Schwenk, Dustin and
Atkinson, David and
Authur, Russell and
Bogin, Ben and
Chandu, Khyathi and
Dumas, Jennifer and
Elazar, Yanai and
Hofmann, Valentin and
Jha, Ananya and
Kumar, Sachin and
Lucy, Li and
Lyu, Xinxi and
Lambert, Nathan and
Magnusson, Ian and
Morrison, Jacob and
Muennighoff, Niklas and
Naik, Aakanksha and
Nam, Crystal and
Peters, Matthew and
Ravichander, Abhilasha and
Richardson, Kyle and
Shen, Zejiang and
Strubell, Emma and
Subramani, Nishant and
Tafjord, Oyvind and
Walsh, Evan and
Zettlemoyer, Luke and
Smith, Noah and
Hajishirzi, Hannaneh and
Beltagy, Iz and
Groeneveld, Dirk and
Dodge, Jesse and
Lo, Kyle",
editor = "Ku, Lun-Wei and
Martins, Andre and
Srikumar, Vivek",
booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = aug,
year = "2024",
address = "Bangkok, Thailand",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.acl-long.840/",
doi = "10.18653/v1/2024.acl-long.840",
pages = "15725--15788"
}