Cross-gender Voice Conversion with Constant F0-Ratio and Average Background Conversion Model
Zbigniew Latka, Jakub Galka, Bartosz Ziólko
Abstract
This paper presents the method for spectral voice conversion using parallel training data. The proposed solution was submitted to the 2018 Voice Conversion Challenge. The method focuses on the preparation of the generative model for cross-gender voice conversion in differential-filtering framework. To improve the quality of the Gaussian mixture conversion model we introduced the usage of the averaged speaker background model pre-training step. Constant F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> ratio transformation of source speech using WORLD vocoder was also proposed to improve cross-gender conversion quality. The evaluation results show that the proposed solution outperforms most of the concurrent systems submitted to the 2018 Voice Conversion Challenge, both in terms of speech quality and similarity. The system achieved 76% similarity score and 3.22 mean opinion score in cross-gender conversion task.
BibTeX
@inproceedings{icassp2019_crossgendervoice,
title = {Cross-gender Voice Conversion with Constant F0-Ratio and Average Background Conversion Model},
author = {Zbigniew Latka and Jakub Galka and Bartosz Ziólko},
booktitle = {ICASSP 2019},
year = {2019}
}