← Search

Steven Moran

3 accepted papers

2024

A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets

NAACL 2024findings

Typologically diverse benchmarks are increasingly created to track the progress achieved in multilingual NLP. Linguistic diversity of these data sets is typically measured as the number of languages or language families included in the sample, but such measures do not consider structural properties…

2024

Converting Legacy Data to CLDF: A FAIR Exit Strategy for Linguistic Web Apps

COLING 2024main

In the mid 2000s, there were several large-scale US National Science Foundation (NSF) grants awarded to projects aiming at developing digital infrastructure and standards for different forms of linguistics data. For example, MultiTree encoded language family trees as phylogenies in XML and LL-MAP co…

Cited by 1SourcePDFScholar
2024

Phonetic Segmentation of the UCLA Phonetics Lab Archive

COLING 2024main

Research in speech technologies and comparative linguistics depends on access to diverse and accessible speech data. The UCLA Phonetics Lab Archive is one of the earliest multilingual speech corpora, with long-form audio recordings and phonetic transcriptions for 314 languages (Ladefoged et al., 200…