ICASSP 2016accepted0 citations

Unsupervised speaker adaptation for DNN-based TTS synthesis

Yuchen Fan, Yao Qian, Frank K. Soong, Lei He

Abstract

Multi-speaker TTS trained with a general DNN has outperformed individually modelled baseline [1]. Multi-speaker DNN takes advantages of larger amount of training data from multiple speakers to find robust transformations in the hidden layers and covers more speaker variability in the output regression layer. In this paper, we propose a new approach to unsupervised speaker adaptation with multi-speaker DNN. It takes advantage of shared hidden transformation to search for the labels of unlabelled acoustic frames and the found labels are used for speaker adaption. Experimental results show that the new approach of unsupervised adaptation can achieve comparable performance with supervised adaptation both objectively and subjectively. We further extend it to cross-lingual adaptation. It can remove non-native accent and improve the naturalness while keep the same speaker's characteristics.

BibTeX
@inproceedings{icassp2016_unsupervisedspea,
  title = {Unsupervised speaker adaptation for DNN-based TTS synthesis},
  author = {Yuchen Fan and Yao Qian and Frank K. Soong and Lei He},
  booktitle = {ICASSP 2016},
  year = {2016}
}