ICASSP 2025accepted0 citations

SPSinger: Multi-Singer Singing Voice Synthesis with Short Reference Prompt

Junchuan Zhao, Chetwin Low, Ye Wang

Abstract

Current singing voice synthesis systems often struggle in multi-singer scenarios due to limited training data that only includes a few singers. Existing zero-shot multi-singer singing voice synthesis systems are criticized for their reliance on global timbre embeddings from single reference audio, which fail to capture sufficient timbre details. This paper introduces SPSinger, a multi-singer singing voice synthesizer that generates singer-specific voices from brief reference audio (around 5 seconds) without prior training on the singer’s voice. SPSinger builds on the StableDiffusion framework by adding a global encoder to capture consistent timbre features from short reference prompts and an attention-based local encoder to capture detailed variations from long prompts, used only during training. To overcome the challenge of requiring long audio prompts during inference, we introduce the Latent Prompt Adaptation Model (LPAM), a Transformer-based module that derives timbre features from global embeddings. This approach eliminates the need for long reference prompts. Additionally, we propose a novel pitch shift algorithm that uses LPAM to predict the pitch shift values. Our experiments show that SPSinger achieves high-quality singing voice synthesis that preserves the identity of the target singer, even when using only short reference audio inputs in zero-shot scenarios.

BibTeX
@inproceedings{icassp2025_spsingermultisin,
  title = {SPSinger: Multi-Singer Singing Voice Synthesis with Short Reference Prompt},
  author = {Junchuan Zhao and Chetwin Low and Ye Wang},
  booktitle = {ICASSP 2025},
  year = {2025}
}
SPSinger: Multi-Singer Singing Voice Synthesis with Short Reference Prompt · ICASSP 2025