Multi-stream spectral representation for statistical parametric speech synthesis
Kayoko Yanagisawa, Ranniery Maia, Yannis Stylianou
Abstract
In statistical parametric speech synthesis such as Hidden Markov Model (HMM) based synthesis, one of the problems is in the over-smoothing of parameters, which leads to a muffled sensation in the synthesised output. In this paper, we propose an approach in which the high frequency spectrum is modelled separately from the low frequency spectrum. The high frequency band, which does not carry much linguistic information, is clustered using a very large decision tree so as to generate parameters as close as possible to natural speech samples. The boundary frequency can be adjusted at synthesis time for each state. Subjective listening tests show that the proposed approach is significantly preferred over the conventional approach using a single spectrum stream. Samples synthesised using the proposed approach sound less muffled and more natural.
BibTeX
@inproceedings{icassp2016_multistreamspect,
title = {Multi-stream spectral representation for statistical parametric speech synthesis},
author = {Kayoko Yanagisawa and Ranniery Maia and Yannis Stylianou},
booktitle = {ICASSP 2016},
year = {2016}
}