ICASSP 2024accepted0 citations

Accent-Specific Vector Quantization for Joint Unsupervised and Supervised Training in Accent Robust Speech Recognition

Li Li, Yijie Li, Dongxing Xu, Haoran Wei, Yanhua Long

Abstract

How to effectively use limited supervised accent data to improve the accented ASR is of paramount importance. In this work, we propose an accent-specific quantization for joint unsupervised and supervised training (AQ-JUST) of end-to-end ASR models to address this issue. Specifically, two variants of AQ-JUST are investigated, namely BAQ-JUST and SAQ-JUST, by employing different model structures and training methods to capture the distinctiveness and commonalities between diverse accents, thus enhancing the performance of accented ASR systems. Our experiments are performed on both accented English and Mandarin ASR tasks. Results show that the proposed methods outperform the strong JUST baseline by relative 3.9% to 9.4% word/character error rate reductions on accented test sets.

BibTeX
@inproceedings{icassp2024_accentspecificve,
  title = {Accent-Specific Vector Quantization for Joint Unsupervised and Supervised Training in Accent Robust Speech Recognition},
  author = {Li Li and Yijie Li and Dongxing Xu and Haoran Wei and Yanhua Long},
  booktitle = {ICASSP 2024},
  year = {2024}
}
Accent-Specific Vector Quantization for Joint Unsupervised and Supervised Training in Accent Robust Speech Recognition · ICASSP 2024