The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data Only
Large language models are commonly trained on a mixture of filtered web data and curated ``high-quality'' corpora, such as social media conversations, books, or technical papers. This curation process is believed to be necessary to produce performant models with broad zero-shot generalization abilit…