ICASSP 2016accepted0 citations
Speaker and language factorization in DNN-based TTS synthesis
Yuchen Fan, Yao Qian, Frank K. Soong, Lei He
Abstract
We have successfully proposed to use multi-speaker modelling in DNN-based TTS synthesis for improved voice quality with limited available data from a speaker. In this paper, we propose a new speaker and language factorized DNN, where speaker-specific layers are used for multi-speaker modelling, and shared layers and language-specific layers are employed for multi-language, linguistic feature transformation. Experimental results on a speech corpus of multiple speakers in both Mandarin and English show that the proposed factorized DNN can not only achieve a similar voice quality as that of a multi-speaker DNN, but also perform polyglot synthesis with a monolingual speaker's voice.
BibTeX
@inproceedings{icassp2016_speakerandlangua,
title = {Speaker and language factorization in DNN-based TTS synthesis},
author = {Yuchen Fan and Yao Qian and Frank K. Soong and Lei He},
booktitle = {ICASSP 2016},
year = {2016}
}