2024
Data Collection Pipeline for Low-Resource Languages: A Case Study on Constructing a Tetun Text Corpus
COLING 2024main
This paper proposes Labadain Crawler, a data collection pipeline tailored to automate and optimize the process of constructing textual corpora from the web, with a specific target to low-resource languages. The system is built on top of Nutch, an open-source web crawler and data extraction framework…