A CURATEd CATalog: Rethinking the Extraction of Pretraining Corpora for Mid-Resourced Languages
We present and describe two language resources in this paper: CATalog 1.0, the largest text corpus in Catalan to date, and CURATE (Corpus Utility for RAting TExt), a modular, parallelizable pipeline used for processing and scoring documents based on text quality that we have optimised to run in High…