A speaker adaptation technique for Gaussian process regression based speech synthesis using feature space transform
Tomoki Koriyama, Syohei Oshio, Takao Kobayashi
Abstract
In this paper, we propose a speaker adaptation technique for statistical parametric speech synthesis based on Gaussian process regression (GPR). Although it is reported that the GPR-based speech synthesis improves the naturalness of synthetic speech compared with the HMM-based speech synthesis, any speaker adaptation techniques for the GPR-based one have not been established. This is because GPR is a nonparametric model and hence it is impossible to directly apply linear transforms to model parameters. In the proposed technique, we introduce feature-space transform to achieve model adaptation in the framework of GPR-based speech synthesis. Experimental results of objective and subjective tests show that the proposed technique outperforms the conventional HMM-based speaker adaptation framework.
BibTeX
@inproceedings{icassp2016_aspeakeradaptati,
title = {A speaker adaptation technique for Gaussian process regression based speech synthesis using feature space transform},
author = {Tomoki Koriyama and Syohei Oshio and Takao Kobayashi},
booktitle = {ICASSP 2016},
year = {2016}
}