← Search

Jorge Palomar-Giner

2 accepted papers

2024

A CURATEd CATalog: Rethinking the Extraction of Pretraining Corpora for Mid-Resourced Languages

COLING 2024main

We present and describe two language resources in this paper: CATalog 1.0, the largest text corpus in Catalan to date, and CURATE (Corpus Utility for RAting TExt), a modular, parallelizable pipeline used for processing and scoring documents based on text quality that we have optimised to run in High…

2024

Building a Data Infrastructure for a Mid-Resource Language: The Case of Catalan

COLING 2024main

Current LLM-based applications are becoming steadily available for everyone with a reliable access to technology and the internet. These applications offer benefits to their users that leave those without access to them at a serious disadvantage. Given the vastly large amount of data needed to train…