← Search

Romi Stella

1 accepted papers

2023

MADLAD-400: A Multilingual And Document-Level Large Audited Dataset

NeurIPS 2023poster

We introduce MADLAD-400, a manually audited, general domain 3T token monolingual dataset based on CommonCrawl, spanning 419 languages. We discuss the limitations revealed by self-auditing MADLAD-400, and the role data auditing had in the dataset creation process. We then train and release a 10.7B-pa…

Cited by 126SourcePDFScholar