Speech Intelligibility Classifiers from 550k Disordered Speech Samples
Subhashini Venugopalan, Jimmy Tobin, Samuel J. Yang, Katie Seaver, Richard J. N. Cave, Pan-Pan Jiang, Neil Zeghidour, Rus Heywood
Abstract
We developed dysarthric speech intelligibility classifiers on 551,176 disordered speech samples contributed by a diverse set of 468 speakers, with a range of self-reported speaking disorders and rated for their overall intelligibility on a five-point scale. We trained three models following different deep learning approaches and evaluated them on ~ 94K utterances from 100 speakers. We further found the models to generalize well (without further training) on the TORGO database[1] (100% accuracy), UASpeech[2] (0.93 correlation), ALS-TDI PMP[3] (0.81 AUC) datasets as well as on a dataset of realistic unprompted speech we gathered (106 dysarthric and 76 control speakers, ~ 2300 samples). We share our model <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> to advance research in this domain.
BibTeX
@inproceedings{icassp2023_speechintelligib,
title = {Speech Intelligibility Classifiers from 550k Disordered Speech Samples},
author = {Subhashini Venugopalan and Jimmy Tobin and Samuel J. Yang and Katie Seaver and Richard J. N. Cave and Pan-Pan Jiang and Neil Zeghidour and Rus Heywood and Jordan R. Green and Michael P. Brenner},
booktitle = {ICASSP 2023},
year = {2023}
}