ICASSP 2022accepted0 citations

HiFi-SVC: Fast High Fidelity Cross-Domain Singing Voice Conversion

Yong Zhou, Xiangju Lu

Abstract

This paper presents HiFi-SVC, a small cross-domain singing voice conversion model for generating high-fidelity 22.05 kHz singing voices. Building on state-of-the-art neural vocoder HiFi-GAN and a convolution-based module for modeling F0, HiFi-SVC can be trained end-to-end with either speech or singing data, achieving better voice similarity on two of the datasets than FastSVC while using slightly smaller number of parameters. We also propose a pitch adjustment method for improving conversion quality.

BibTeX
@inproceedings{icassp2022_hifisvcfasthighf,
  title = {HiFi-SVC: Fast High Fidelity Cross-Domain Singing Voice Conversion},
  author = {Yong Zhou and Xiangju Lu},
  booktitle = {ICASSP 2022},
  year = {2022}
}
HiFi-SVC: Fast High Fidelity Cross-Domain Singing Voice Conversion · ICASSP 2022