A comparative study of acoustic-to-articulatory inversion for neutral and whispered speech
Aravind Illa, Nisha Meenakshi, Prasanta Kumar Ghosh
Abstract
Whispered speech is known to have different characteristics in acoustics and articulation compared to neutral speech. In this study, we compare the accuracy with which the articulation can be recovered from the acoustics of both types of speech, individually. Acoustic-to-articulatory inversion (AAI) is performed with twelve articulatory features using the deep neural network (DNN) with data obtained from four subjects. We consider AAI in matched and mis-matched train-test conditions, where the speech types in training and test are identical and different respectively. Experiments in matched condition reveal that the AAI performance for whispered speech drops significantly compared to that for neutral speech, only for jaw, tongue tip and tongue body, consistently, for all four subjects. This indicates that the whispered speech encodes information about the rest of the articulators to a degree similar to that of the neutral speech. Experiments in the mis-matched condition show a consistent drop in the AAI performance compared to the matched condition. This drop in performance from matched to mis-matched condition is found be the highest for upper lip which indicates that the upper lip movement could be encoded differently in whispered speech compared to that in neutral speech.
BibTeX
@inproceedings{icassp2017_acomparativestud,
title = {A comparative study of acoustic-to-articulatory inversion for neutral and whispered speech},
author = {Aravind Illa and Nisha Meenakshi and Prasanta Kumar Ghosh},
booktitle = {ICASSP 2017},
year = {2017}
}