ICASSP 2025accepted0 citations

Dysarthric Speech Conformer: Adaptation for Sequence-to-Sequence Dysarthric Speech Recognition

Qianli Wang, Zihan Zhong, Satwinder Singh, Clarion Mendes, Mark Hasegawa-Johnson, Waleed Abdulla, Seyed Reza Shahamiri

Abstract

Automatic Speech Recognition (ASR) holds immense potential to provide an effective interface for assistive technologies, but its performance remains unsatisfactory for people with speech impairments such as dysarthria. Existing ASR systems struggle to accurately recognize dysarthric speech due to the significant speaker variability in dysarthric speech and the scarcity of dysarthric datasets. In this study, we propose a two-phase adaptation pipeline based on the Conformer architecture that leverages typical speech to transfer to individualized ASR models for dysarthric speakers. ASR performance is evaluated for isolated words and continuous sentences, yielding an average Word Error Rate of 21.5% on the UASpeech dataset and 12.7% on the TORGO dataset. Selectively freezing decoder layers was more often successful than selectively freezing encoder layers, suggesting that optimal performance is achieved by focusing the adaptation on the acoustic information contained in the encoder.

BibTeX
@inproceedings{icassp2025_dysarthricspeech,
  title = {Dysarthric Speech Conformer: Adaptation for Sequence-to-Sequence Dysarthric Speech Recognition},
  author = {Qianli Wang and Zihan Zhong and Satwinder Singh and Clarion Mendes and Mark Hasegawa-Johnson and Waleed Abdulla and Seyed Reza Shahamiri},
  booktitle = {ICASSP 2025},
  year = {2025}
}