From Characters to Subwords: Modeling Unit Conversion for Low-resource Speech Recognition
Yizhi Wang, Haofei Zhang, Huiqiong Wang, Li Sun, Mingli Song
Abstract
Multilingual automatic speech recognition (ASR) models greatly facilitate recognizing low-resource languages by sharing representations across similar languages. However, the commonly adopted modeling units, e.g., character-level modeling, lack language-specific information, resulting in a susceptible word prediction to phonemes and characters. Recently, subword-level modeling has demonstrated significant effectiveness for monolingual automatic recognition systems, while it is adverse to cross-lingual feature sharing. In this paper, we propose a novel low-resource ASR method that leverages the advantages of two different modeling units. Specifically, a character-level ASR model is trained on the multilingual dataset for modeling the short-term speech and learning general speech knowledge from relevant languages. Afterwards, we convert the character-level prediction into subwords for learning contextual information of the target language. Extensive experiments on Uyghur with Kazakh and Kyrgyz as auxiliary languages have shown that our proposed method significantly reduces word error rate (WER).
BibTeX
@inproceedings{icassp2025_fromcharactersto,
title = {From Characters to Subwords: Modeling Unit Conversion for Low-resource Speech Recognition},
author = {Yizhi Wang and Haofei Zhang and Huiqiong Wang and Li Sun and Mingli Song},
booktitle = {ICASSP 2025},
year = {2025}
}