ICASSP 2016accepted0 citations

Robust TTS duration modelling using DNNS

Gustav Eje Henter, Srikanth Ronanki, Oliver Watts, Mirjam Wester, Zhizheng Wu, Simon King

Abstract

Accurate modelling and prediction of speech-sound durations is an important component in generating more natural synthetic speech. Deep neural networks (DNNs) offer a powerful modelling paradigm, and large, found corpora of natural and expressive speech are easy to acquire for training them. Unfortunately, found datasets are seldom subject to the quality-control that traditional synthesis methods expect. Common issues likely to affect duration modelling include transcription errors, reductions, filled pauses, and forced-alignment inaccuracies. To combat this, we propose to improve modelling and prediction of speech durations using methods from robust statistics, which are able to disregard ill-fitting points in the training material. We describe a robust fitting criterion based on the density power divergence (the ß-divergence) and a robust generation heuristic using mixture density networks (MDNs). Perceptual tests indicate that subjects prefer synthetic speech generated using robust models of duration over the baselines.

BibTeX
@inproceedings{icassp2016_robustttsduratio,
  title = {Robust TTS duration modelling using DNNS},
  author = {Gustav Eje Henter and Srikanth Ronanki and Oliver Watts and Mirjam Wester and Zhizheng Wu and Simon King},
  booktitle = {ICASSP 2016},
  year = {2016}
}
Robust TTS duration modelling using DNNS · ICASSP 2016