Integrating Potential Pronunciations for Enhanced Mispronunciation Detection and Diagnosis Ability in LLMs
Minglin Wu, Jing Xu, Xueyuan Chen, Helen Meng
Abstract
Large Language Models (LLMs) have exhibited significant potentials across various tasks. However, how to leverage the power of LLMs in the mispronunciation detection and diagnosis (MDD) task is still under-explored. In this paper, we propose a PP-ATP model, which integrates potential pronunciations covering common mispronunciations into the prompt part of LLMs, to enhance the MDD ability of LLMs in second language (L2) English. Specifically, the proposed PP-ATP model is composed of an audio encoder, an LLM decoder, and an adapter. Taking speech representations from the audio encoder as the audio prompt and reference sentence with canonical and potential pronunciations as text prompt, the LLM decoder is adapted to predict the actual pronunciation in the given L2 speech. Experiments show that our PP-ATP model achieves new state-of-the-art (SOTA) performance in MDD on CU-CHLOE corpus, confirming the effectiveness of potential pronunciation integration.
BibTeX
@inproceedings{icassp2025_integratingpoten,
title = {Integrating Potential Pronunciations for Enhanced Mispronunciation Detection and Diagnosis Ability in LLMs},
author = {Minglin Wu and Jing Xu and Xueyuan Chen and Helen Meng},
booktitle = {ICASSP 2025},
year = {2025}
}