ICASSP 2016accepted0 citations

Directly modeling voiced and unvoiced components in speech waveforms by neural networks

Keiichi Tokuda, Heiga Zen

Abstract

This paper proposes a novel acoustic model based on neural networks for statistical parametric speech synthesis. The neural network outputs parameters of a non-zero mean Gaussian process, which defines a probability density function of a speech waveform given linguistic features. The mean and covariance functions of the Gaussian process represent deterministic (voiced) and stochastic (unvoiced) components of a speech waveform, whereas the previous approach considered the unvoiced component only. Experimental results show that the proposed approach can generate speech waveforms approximating natural speech waveforms.

BibTeX
@inproceedings{icassp2016_directlymodeling,
  title = {Directly modeling voiced and unvoiced components in speech waveforms by neural networks},
  author = {Keiichi Tokuda and Heiga Zen},
  booktitle = {ICASSP 2016},
  year = {2016}
}
Directly modeling voiced and unvoiced components in speech waveforms by neural networks · ICASSP 2016