COLING 2025main0 citations

A Benchmark of French ASR Systems Based on Error Severity

Antoine Tholly, Jane Wottawa, Mickael Rouvier, Richard Dufour

Abstract

Automatic Speech Recognition (ASR) transcription errors are commonly assessed using metrics that compare them with a reference transcription, such as Word Error Rate (WER), which measures spelling deviations from the reference, or semantic score-based metrics. However, these approaches often overlook what is understandable to humans when interpreting transcription errors. To address this limitation, a new evaluation is proposed that categorizes errors into four levels of severity, further divided into subtypes, based on objective linguistic criteria, contextual patterns, and the use of content words as the unit of analysis. This metric is applied to a benchmark of 10 state-of-the-art ASR systems on French language, encompassing both HMM-based and end-to-end models. Our findings reveal the strengths and weaknesses of each system, identifying those that provide the most comfortable reading experience for users.

BibTeX
@inproceedings{tholly-etal-2025-benchmark,
    title = "A Benchmark of {F}rench {ASR} Systems Based on Error Severity",
    author = "Tholly, Antoine  and
      Wottawa, Jane  and
      Rouvier, Mickael  and
      Dufour, Richard",
    editor = "Rambow, Owen  and
      Wanner, Leo  and
      Apidianaki, Marianna  and
      Al-Khalifa, Hend  and
      Eugenio, Barbara Di  and
      Schockaert, Steven",
    booktitle = "Proceedings of the 31st International Conference on Computational Linguistics",
    month = jan,
    year = "2025",
    address = "Abu Dhabi, UAE",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.coling-main.341/",
    pages = "5094--5101"
}