ICASSP 2025accepted0 citations

Unveiling Performance Bias in ASR Systems: A Study on Gender, Age, Accent, and More

Maliha Jahan, Priyam Mazumdar, Thomas Thebaud, Mark Hasegawa-Johnson, Jesús Villalba, Najim Dehak, Laureano Moro-Velázquez

Abstract

With the recent advancements in speech recognition, it is crucial to ensure these systems are free from performance biases against any speaker subgroups. This study examined the performance of twenty variants of seven Automatic Speech Recognition models across four datasets in English language: L2 Arctic, Speech Accent Archive, CORAAL, and SBCSAE. We employed Poisson regression and drop-in-deviance tests to identify which attributes significantly contribute to the Word Error Rate. Our analysis revealed biases related to attributes such as native language, location, occupation, and birthplace. Most systems did not exhibit bias related to factors like gender and age. Additionally, we conducted an experiment to detect bias related to "variant" (accent and dialect) by combining the CORAAL (African American Vernacular English (AAVE)) and SBCSAE (General American English (GAE)) datasets, aiming to identify the sources of any observed bias. We found that both speaker variability and dialectal difference contribute to observed bias for variant.

BibTeX
@inproceedings{icassp2025_unveilingperform,
  title = {Unveiling Performance Bias in ASR Systems: A Study on Gender, Age, Accent, and More},
  author = {Maliha Jahan and Priyam Mazumdar and Thomas Thebaud and Mark Hasegawa-Johnson and Jesús Villalba and Najim Dehak and Laureano Moro-Velázquez},
  booktitle = {ICASSP 2025},
  year = {2025}
}