← Search

Anton Lozhkov

2 accepted papers

2024

The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

NeurIPS 2024spotlight

The performance of a large language model (LLM) depends heavily on the quality and size of its pretraining dataset. However, the pretraining datasets for state-of-the-art open LLMs like Llama 3 and Mixtral are not publicly available and very little is known about how they were created. In this work,…

Cited by 86SourcePDFScholar
2023

OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

NeurIPS 2023poster

Large multimodal models trained on natural documents, which interleave images and text, outperform models trained on image-text pairs on various multimodal benchmarks. However, the datasets used to train these models have not been released, and the collection process has not been fully specified. W…