ICASSP 2025accepted0 citations

PriorSinger: Singing Voice Synthesis Model with Prior Condition Cross Attention

Zehua Zhang, Bosong Yan, Yinghan Cao, Mingjiang Wang

Abstract

The singing voice synthesis system is designed to generate realistic and expressive singing based on a given musical score. Generative Adversarial Networks (GANs) or diffusion models generate acoustic features, such as Mel-spectrograms, which are subsequently reconstructed into waveforms by a vocoder. In this work, the musical score is encoded as a prior condition to guide the diffusion denoiser through a novel prior cross-attention Transformer during the denoising process. Moreover, we introduce attention mechanisms in both the time and frequency domains within the diffusion denoiser to enhance the resolution of the generated acoustic features. Additionally, incorporating rotary positional encoding allows the model to better handle temporal and frequency positional information. Our model is capable of synthesizing singing with both higher quality and more vivid expressiveness. In subjective evaluations of the Opencpop dataset, our model outperforms state-of-the-art methods.

BibTeX
@inproceedings{icassp2025_priorsingersingi,
  title = {PriorSinger: Singing Voice Synthesis Model with Prior Condition Cross Attention},
  author = {Zehua Zhang and Bosong Yan and Yinghan Cao and Mingjiang Wang},
  booktitle = {ICASSP 2025},
  year = {2025}
}