ICASSP 2016accepted0 citations

Implementation of F0 transformation for statistical singing voice conversion based on direct waveform modification

Kazuhiro Kobayashi, Tomoki Toda, Satoshi Nakamura

Abstract

This paper presents a technique for transforming F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> in a framework of statistical singing voice conversion with direct waveform modification based on spectrum differential (DIFFSVC). The DIFFSVC method converts voice timbre of singing voices of a source singer into that of a target singer without using vocoder-based waveform generation. Although this method achieves high sound quality of the converted singing voices, its use is limited to only intra-gender conversion without the need of F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> transformation. To make it possible to also use the DIFFSVC method for cross-gender conversion, we propose a method to transform F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> of an input singing voice for the DIFFSVC. The proposed method is also based on direct waveform modification using overlap-add process and filtering process. Results of subjective evaluations demonstrate that the proposed DIFFSVC method with F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> transformation significantly improves sound quality of the converted singing voices while preserving the conversion accuracy of singer identity in the cross-gender conversion compared to the conventional SVC with vocoder.

BibTeX
@inproceedings{icassp2016_implementationof,
  title = {Implementation of F0 transformation for statistical singing voice conversion based on direct waveform modification},
  author = {Kazuhiro Kobayashi and Tomoki Toda and Satoshi Nakamura},
  booktitle = {ICASSP 2016},
  year = {2016}
}