ICASSP 2019accepted0 citations

CNN-RNN-CTC Based End-to-end Mispronunciation Detection and Diagnosis

Wai-Kim Leung, Xunying Liu, Helen Meng

Abstract

This paper focuses on using Convolutional Neural Network (CNN), Recurrent Neural Network (RNN) and Connection-ist Temporal Classification (CTC) to build an end-to-end speech recognition for Mispronunciation Detection and Diagnosis (MDD) task. Our approach is end-to-end models, while phonemic or graphemic information, or forced alignment between different linguistic units, are not required. We conduct experiments that compare the proposed CNN-RNN-CTC approach with alternative mispronunciation detection and diagnoses (MDD) approaches. The F-measure of our approach is 74.65%, which significantly outperforms the Extended Recognition Network (ERN) (S-AM) by 44.75% and State-level Acoustic Model (S-AM) by 32.28% relatively. The relative improvement in F-measure when over Acoustic-Phonemic Model (APM), Acoustic-Graphemic Model (AGM) and Acoustic-Phonemic-Graphemic Model (APGM) are 9.57%, 5.04% and 2.77% respectively.

BibTeX
@inproceedings{icassp2019_cnnrnnctcbaseden,
  title = {CNN-RNN-CTC Based End-to-end Mispronunciation Detection and Diagnosis},
  author = {Wai-Kim Leung and Xunying Liu and Helen Meng},
  booktitle = {ICASSP 2019},
  year = {2019}
}