ICASSP 2025accepted0 citations

FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion

Alef Iury Siqueira Ferreira, Lucas Rafael Stefanel Gris, Augusto Seben da Rosa, Frederico Santos de Oliveira, Edresson Casanova, Rafael Teixeira Sousa, Arnaldo Cândido Jr., Anderson da Silva Soares

Abstract

This work presents FreeSVC, a promising multilingual singing voice conversion approach that leverages an enhanced VITS model with Speaker-invariant Clustering (SPIN) for better content representation and the State-of-the-Art (SOTA) speaker encoder ECAPA2. FreeSVC incorporates trainable language embeddings to handle multiple languages and employs an advanced speaker encoder to disentangle speaker characteristics from linguistic content. Designed for zero-shot learning, FreeSVC enables cross-lingual singing voice conversion without extensive language-specific training. We demonstrate that a multilingual content extractor is crucial for optimal cross-language conversion. Our source code and models are publicly available<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>.

BibTeX
@inproceedings{icassp2025_freesvctowardsze,
  title = {FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion},
  author = {Alef Iury Siqueira Ferreira and Lucas Rafael Stefanel Gris and Augusto Seben da Rosa and Frederico Santos de Oliveira and Edresson Casanova and Rafael Teixeira Sousa and Arnaldo Cândido Jr. and Anderson da Silva Soares and Arlindo R. Galvão Filho},
  booktitle = {ICASSP 2025},
  year = {2025}
}