EMNLP 20250 citations

Identifying Rare Languages in Common Crawl Data is a Needles-in-a-Haystack Problem

Rasul Dent, Pedro Ortiz Suarez, Thibault Cl{\'e}rice, Beno{\^i}t Sagot

Abstract

Automatic language identification is frequently framed as a multi-class classification problem. However, when creating digital corpora for less commonly written languages, it may be more appropriate to consider it a data mining problem. For these varieties, one knows ahead of time that the vast majority of documents are of little interest. By minimizing resources spent on classifying such documents, we can create corpora covering previously overlooked languages faster than existing pipelines. To demonstrate the effectiveness of the targeted mining perspective, we introduce a new pipeline that can filter a single snapshot in two hours. We also provide web corpora for several French-based Creoles.

BibTeX
@inproceedings{emnlp2025_identifyingrarel,
  title = {Identifying Rare Languages in Common Crawl Data is a Needles-in-a-Haystack Problem},
  author = {Rasul Dent and Pedro Ortiz Suarez and Thibault Cl{\'e}rice and Beno{\^i}t Sagot},
  booktitle = {EMNLP 2025},
  year = {2025}
}
Identifying Rare Languages in Common Crawl Data is a Needles-in-a-Haystack Problem · EMNLP 2025