ICASSP 2020accepted0 citations

Semi-Supervised Learning Based on Hierarchical Generative Models for End-to-End Speech Synthesis

Takato Fujimoto, Shinji Takaki, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda

Abstract

This paper proposes a general framework of semi-supervised learning based on hierarchical generative models and adapts it to a Japanese end-to-end text-to-speech (TTS) system. In English TTS, several end-to-end systems have recently achieved sound quality close to that of natural human speech. However, in non-alphabetic languages such as Japanese, it is difficult to realize true text-input end-to-end TTS due to character diversity and pitch accents. To address this problem, we propose end-to-end TTS based on semi-supervised learning that makes the most of existing data consisting of any combination of text, phoneme, and waveform as training data. To demonstrate the effectiveness of the proposed system, listening tests were conducted for pronunciation and naturalness. Our results show that the proposed system improves both pronunciation and naturalness.

BibTeX
@inproceedings{icassp2020_semisupervisedle,
  title = {Semi-Supervised Learning Based on Hierarchical Generative Models for End-to-End Speech Synthesis},
  author = {Takato Fujimoto and Shinji Takaki and Kei Hashimoto and Keiichiro Oura and Yoshihiko Nankaku and Keiichi Tokuda},
  booktitle = {ICASSP 2020},
  year = {2020}
}
Semi-Supervised Learning Based on Hierarchical Generative Models for End-to-End Speech Synthesis · ICASSP 2020