ICASSP 2018accepted0 citations

Vae-Space: Deep Generative Model of Voice Fundamental Frequency Contours

Kou Tanaka, Hirokazu Kameoka, Kazuho Morikawa

Abstract

Modeling the speech generation process can provide flexible and interpretable ways to generate intended synthetic speech. In this paper, we present a deep generative model of fundamental frequency (F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> ) contours of normal speech and singing voices. The generative model we propose in this paper 1) is able to accurately decompose an F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> contour into the sum of phrase and accent components of the Fujisaki model, a mathematical model describing the control mechanism of vocal fold vibration, without an iterative algorithm, and 2) can represent/generate F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> contours of both normal speech and singing voices reasonably well.

BibTeX
@inproceedings{icassp2018_vaespacedeepgene,
  title = {Vae-Space: Deep Generative Model of Voice Fundamental Frequency Contours},
  author = {Kou Tanaka and Hirokazu Kameoka and Kazuho Morikawa},
  booktitle = {ICASSP 2018},
  year = {2018}
}