Identifying Rare Languages in Common Crawl Data is a Needles-in-a-Haystack Problem
Automatic language identification is frequently framed as a multi-class classification problem. However, when creating digital corpora for less commonly written languages, it may be more appropriate to consider it a data mining problem. For these varieties, one knows ahead of time that the vast majo