ICASSP 2025accepted0 citations

Weak-to-Strong Generalization in Speech Recognition

Soheil Khorram, Qian Zhang, Rohit Prabhavalkar, Kartik Audhkhasi, Bhuvana Ramabhadran

Abstract

To surpass human-level accuracy, speech recognition models must go beyond relying solely on human labels. To this end, we must build stronger models from weaker supervisors and this is the main goal in weak-to-strong generalization (WSG). WSG methods normally incorporate additional information into weak teacher models to improve their performance for example reliability of teacher-generated labels. In this research, we investigate two sources of additional information to implement WSG for speech recognition: unsupervised data from the target language and supervised data from other similar languages. We study scenarios where unsupervised data boosts performance, and propose a new clustering method leveraging supervised data from similar languages for further gains. Our clustering method yields an average 8% reduction in word error rate compared to a universal speech model trained on 182 languages. By incorporating both supervised data from similar languages and unsupervised data from the target language, we further enhance the USM model by 10%. This improvement reaches 15% for the top-performing languages with a WER below 50%.

BibTeX
@inproceedings{icassp2025_weaktostronggene,
  title = {Weak-to-Strong Generalization in Speech Recognition},
  author = {Soheil Khorram and Qian Zhang and Rohit Prabhavalkar and Kartik Audhkhasi and Bhuvana Ramabhadran},
  booktitle = {ICASSP 2025},
  year = {2025}
}