2020
Semi-Supervised Speaker Adaptation for End-to-End Speech Synthesis with Pretrained Models
ICASSP 2020accepted
Recently, end-to-end text-to-speech (TTS) models have achieved a remarkable performance, however, requiring a large amount of paired text and speech data for training. On the other hand, we can easily collect unpaired dozen minutes of speech recordings for a target speaker without corresponding text…