Exemplar-based sparse representation of timbre and prosody for voice conversion
Huaiping Ming, Dong-Yan Huang, Lei Xie, Shaofei Zhang, Minghui Dong, Haizhou Li
Abstract
Voice conversion (VC) aims to make one speaker (source) to sound like spoken by another speaker (target) without changing the language content. Most of the state-of-the-art voice conversion systems focus only on timbre conversion. However, the speaker identity is characterized by the source-related cues such as fundamental frequency and energy as well. In this work, we propose an exemplarbased sparse representation of timbre and prosody for voice conversion that does not necessitate separately timbre conversion and prosody conversions. The experiment results show that, in addition to the conversion of spectral features, the proper conversion of prosody features will improve the quality and speaker identity of the converted speech.
BibTeX
@inproceedings{icassp2016_exemplarbasedspa,
title = {Exemplar-based sparse representation of timbre and prosody for voice conversion},
author = {Huaiping Ming and Dong-Yan Huang and Lei Xie and Shaofei Zhang and Minghui Dong and Haizhou Li},
booktitle = {ICASSP 2016},
year = {2016}
}