Speech Few-Shot Learning for Language Learners' Speech Recognition
Abstract
This paper reports how speech recognition accuracy can be improved using the speech few-shot in-context learning capabilities of a multimodal foundation model when applied to the speech of language learners. Our proposed method, which combines speech few-shot with context prompting, demonstrates significant improvements in recognizing language learners’ accented speech. Evaluations on data from an L2 English test set with accented speech produced a 33.1% relative WER reduction and improved target word recall from 89% to 97% compared to the Gemini 1.5 Pro baseline. Notably, speech few-shot alone contributes a 9.8% relative WER reduction beyond the gains from context prompting. These results underscore the importance of incorporating domain knowledge from both speech and text modalities within in-context learning, suggesting that speech few-shot in-context learning offers a viable alternative to resource-intensive fine-tuning for addressing challenges in atypical speech recognition.
BibTeX
@inproceedings{icassp2025_speechfewshotlea,
title = {Speech Few-Shot Learning for Language Learners' Speech Recognition},
author = {Jian Cheng and Sam Nguyen},
booktitle = {ICASSP 2025},
year = {2025}
}