← Search

Daan van Esch

4 accepted papers

2024

Connecting Language Technologies with Rich, Diverse Data Sources Covering Thousands of Languages

COLING 2024main

Contrary to common belief, there are rich and diverse data sources available for many thousands of languages, which can be used to develop technologies for these languages. In this paper, we provide an overview of some of the major online data sources, the types of data that they provide access to,…

Cited by 0SourcePDFScholar
2024

LinguaMeta: Unified Metadata for Thousands of Languages

COLING 2024main

We introduce LinguaMeta, a unified resource for language metadata for thousands of languages, including language codes, names, number of speakers, writing systems, countries, official status, coordinates, and language varieties. The resources are drawn from various existing repositories and suppleme…

Cited by 0SourcePDFScholar
2024

Multimodal Modeling for Spoken Language Identification

ICASSP 2024accepted

Spoken language identification refers to the task of automatically predicting the spoken language in a given utterance. Conventionally, it is modeled as a speech-based language identification task. Prior techniques have been constrained to a single modality; however in the case of video data there i…

Cited by 0SourceScholar
2020

Language ID in the Wild: Unexpected Challenges on the Path to a Thousand-Language Web Text Corpus

COLING 2020main

Large text corpora are increasingly important for a wide variety of Natural Language Processing (NLP) tasks, and automatic language identification (LangID) is a core technology needed to collect such datasets in a multilingual context. LangID is largely treated as solved in the literature, with mode…