2021
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
EMNLP 2021main
Large language models have led to remarkable progress on many NLP tasks, and researchers are turning to ever-larger text corpora to train them. Some of the largest corpora available are made by scraping significant portions of the internet, and are frequently introduced with only minimal documentati…