ICASSP 2023accepted0 citations

Any-to-Any Voice Conversion with F0 and Timbre Disentanglement and Novel Timbre Conditioning

Sudheer Kovela, Rafael Valle, Ambrish Dantrey, Bryan Catanzaro

Abstract

Despite recent advances in voice conversion (VC), it is still challenging to do real-time one-shot voice conversion with good control over timbre and F <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</inf> . In this work, we present a PPG-based VC model that directly decodes waveforms. We designed a speaker conditioned decoder based on HiFi-GAN[1], along with a new discriminator that produces high quality audio. Using an F <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</inf> prenet and F <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</inf> augmented speaker encoder, we are able to control F <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</inf> and timbre independently with high fidelity. Our objective and subjective evaluations show that our method is preferred over others in terms of audio quality, timbre similarity and prosody retention.

BibTeX
@inproceedings{icassp2023_anytoanyvoicecon,
  title = {Any-to-Any Voice Conversion with F0 and Timbre Disentanglement and Novel Timbre Conditioning},
  author = {Sudheer Kovela and Rafael Valle and Ambrish Dantrey and Bryan Catanzaro},
  booktitle = {ICASSP 2023},
  year = {2023}
}
Any-to-Any Voice Conversion with F0 and Timbre Disentanglement and Novel Timbre Conditioning · ICASSP 2023