ICASSP 2018accepted0 citations

Concatenative Articulatory Video Synthesis Using Real-Time MRI Data for Spoken Language Training

Urvish Desai, Chiranjeevi Yarra, Prasanta Kumar Ghosh

Abstract

Spoken language training benefits from showing a video of native speakers' articulatory movements to train the second language learners. Typically, the articulatory video is prepared in conjunction with the audio which is collected simultaneously with the articulatory recording. Articulatory video recording requires specialized equipment and, hence, is expensive and time consuming. In this work, we propose a concatenative synthesis approach to obtain articulatory videos for an audio, which may not have a simultaneous articulatory recording. In the training stage of the proposed approach, we make a repository for phoneme specific articulatory image sequence from the available articulatory video. During testing, image sequences are selected from this repository to ensure a smooth transition across phonetic events. The selected image sequences are finally stitched to synthesize the articulatory video for the test audio. Articulatory videos are synthesized for 50 words randomly selected from the MRI-TIMIT database, not seen in the training data. Subjective evaluation on the quality of the synthesized videos using twelve subjects suggests that the videos are close to the original ones with a rating of 3.78 out of 5, where a score of 5 (1) indicates that there is no (great) difference in quality between the original and the synthesized videos.

BibTeX
@inproceedings{icassp2018_concatenativeart,
  title = {Concatenative Articulatory Video Synthesis Using Real-Time MRI Data for Spoken Language Training},
  author = {Urvish Desai and Chiranjeevi Yarra and Prasanta Kumar Ghosh},
  booktitle = {ICASSP 2018},
  year = {2018}
}