← Search

Eric Le Ferrand

5 accepted papers

2025

That doesn’t sound right: Evaluating speech transcription quality in field linguistics corpora

ACL 2025short

Incorporating automatic speech recognition (ASR) into field linguistics workflows for language documentation has become increasingly common. While ASR performance has seen improvements in low-resource settings, obstacles remain when training models on data collected by documentary linguists. One not…

2024

Are modern neural ASR architectures robust for polysynthetic languages?

EMNLP 2024finding

Automatic speech recognition (ASR) technology is frequently proposed as a means of preservation and documentation of endangered languages, with promising results thus far. Among the endangered languages spoken today, a significant number exhibit complex morphology. The models employed in contemporar…

2024

How Important is a Language Model for Low-resource ASR?

ACL 2024findings

N-gram language models (LMs) are the innovation that first made large-vocabulary continuous automatic speech recognition (ASR) viable. With neural end-to-end ASR architectures, however, LMs have become an afterthought. While the effect on accuracy may be negligible for English and Mandarin, jettison…

Cited by 1SourcePDFScholar
2022

Learning From Failure: Data Capture in an Australian Aboriginal Community

ACL 2022long

Most low resource language technology development is premised on the need to collect data for training statistical models. When we follow the typical process of recording and transcribing text for small Indigenous languages, we hit up against the so-called “transcription bottleneck.” Therefore it is…

Cited by 13SourcePDFScholar
2020

Enabling Interactive Transcription in an Indigenous Community

COLING 2020main

We propose a novel transcription workflow which combines spoken term detection and human-in-the-loop, together with a pilot experiment. This work is grounded in an almost zero-resource scenario where only a few terms have so far been identified, involving two endangered languages. We show that in th…