ICASSP 2016accepted0 citations

Speaker and language factorization in DNN-based TTS synthesis

Yuchen Fan, Yao Qian, Frank K. Soong, Lei He

Abstract

We have successfully proposed to use multi-speaker modelling in DNN-based TTS synthesis for improved voice quality with limited available data from a speaker. In this paper, we propose a new speaker and language factorized DNN, where speaker-specific layers are used for multi-speaker modelling, and shared layers and language-specific layers are employed for multi-language, linguistic feature transformation. Experimental results on a speech corpus of multiple speakers in both Mandarin and English show that the proposed factorized DNN can not only achieve a similar voice quality as that of a multi-speaker DNN, but also perform polyglot synthesis with a monolingual speaker's voice.

BibTeX
@inproceedings{icassp2016_speakerandlangua,
  title = {Speaker and language factorization in DNN-based TTS synthesis},
  author = {Yuchen Fan and Yao Qian and Frank K. Soong and Lei He},
  booktitle = {ICASSP 2016},
  year = {2016}
}