← Search

Pedro Ortiz Suarez

6 accepted papers

2025

Identifying Rare Languages in Common Crawl Data is a Needles-in-a-Haystack Problem

EMNLP 2025

Automatic language identification is frequently framed as a multi-class classification problem. However, when creating digital corpora for less commonly written languages, it may be more appropriate to consider it a data mining problem. For these varieties, one knows ahead of time that the vast majo

2025

mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus

ACL 2025finding

Multimodal Large Language Models (mLLMs) are trained on a large amount of text-image data. While most mLLMs are trained on caption-like data only, Alayrac et al. (2022) showed that additionally training them on interleaved sequences of text and images can lead to the emergence of in-context learning…

2024

A CURATEd CATalog: Rethinking the Extraction of Pretraining Corpora for Mid-Resourced Languages

COLING 2024main

We present and describe two language resources in this paper: CATalog 1.0, the largest text corpus in Catalan to date, and CURATE (Corpus Utility for RAting TExt), a modular, parallelizable pipeline used for processing and scoring documents based on text quality that we have optimised to run in High…

2024

Tokenizer Choice For LLM Training: Negligible or Crucial?

NAACL 2024findings

The recent success of large language models (LLMs) has been predominantly driven by curating the training dataset composition, scaling of model architectures and dataset sizes and advancements in pretraining objectives, leaving tokenizer influence as a blind spot.Shedding light on this underexplored…

2022

A Data-driven Approach to Named Entity Recognition for Early Modern French

COLING 2022main

Named entity recognition has become an increasingly useful tool for digital humanities research, specially when it comes to historical texts. However, historical texts pose a wide range of challenges to both named entity recognition and natural language processing in general that are still difficult…

2022

The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset

NeurIPS 2022accept

As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop, a 1-year international and multidisciplinary initiative, was formed with the goal of researching and training large lan…

Cited by 214SourcePDFScholar