F0 Estimation for DNN-Based Ultrasound Silent Speech Interfaces
Tamás Grósz, Gábor Gosztolya, László Tóth, Tamás Gábor Csapó, Alexandra Markó
Abstract
State-of-the-art silent speech interface systems apply vocoders to generate the speech signal directly from articulatory data. Most of these approaches concentrate on estimating just the spectral features of the vocoder, and use the original F0, a constant F0 or white noise as excitation. This solution is based on the assumption that the F0 curve is unpredictable from articulatory data that does not contain direct measurements of the vocal fold vibration. Here, we experimented with deep neural networks to perform articulatory-to-acoustic conversion from ultrasound images, with an emphasis on estimating the voicing feature and the F0 curve from the ultrasound input. Contrary to the common belief that F0 is unpredictable, we attained a correlation rate of 0.74 between the original and the predicted F0 curve. What is more, the listening tests revealed that our subjects could not distinguish the sentences synthesized using the DNN-estimated and the original F0 curve, and ranked them as having the same quality.
BibTeX
@inproceedings{icassp2018_f0estimationford,
title = {F0 Estimation for DNN-Based Ultrasound Silent Speech Interfaces},
author = {Tamás Grósz and Gábor Gosztolya and László Tóth and Tamás Gábor Csapó and Alexandra Markó},
booktitle = {ICASSP 2018},
year = {2018}
}