Read to Hear: A Zero-Shot Pronunciation Assessment Using Textual Descriptions and LLMs
Automatic pronunciation assessment is typically performed by acoustic models trained on audio-score pairs. Although effective, these systems provide only numerical scores, without the information needed to help learners understand their errors. Meanwhile, large language models (LLMs) have proven eff