Accent-Specific Vector Quantization for Joint Unsupervised and Supervised Training in Accent Robust Speech Recognition
Li Li, Yijie Li, Dongxing Xu, Haoran Wei, Yanhua Long
Abstract
How to effectively use limited supervised accent data to improve the accented ASR is of paramount importance. In this work, we propose an accent-specific quantization for joint unsupervised and supervised training (AQ-JUST) of end-to-end ASR models to address this issue. Specifically, two variants of AQ-JUST are investigated, namely BAQ-JUST and SAQ-JUST, by employing different model structures and training methods to capture the distinctiveness and commonalities between diverse accents, thus enhancing the performance of accented ASR systems. Our experiments are performed on both accented English and Mandarin ASR tasks. Results show that the proposed methods outperform the strong JUST baseline by relative 3.9% to 9.4% word/character error rate reductions on accented test sets.
BibTeX
@inproceedings{icassp2024_accentspecificve,
title = {Accent-Specific Vector Quantization for Joint Unsupervised and Supervised Training in Accent Robust Speech Recognition},
author = {Li Li and Yijie Li and Dongxing Xu and Haoran Wei and Yanhua Long},
booktitle = {ICASSP 2024},
year = {2024}
}