ICASSP 2025accepted0 citations

Self-supervised Prosody Learning at Phoneme-level with Momentum Contrast for Speech Synthesis

Zhaoci Liu, Ya-Jun Hu, Liping Chen, Zhen-Hua Ling

Abstract

This paper investigates leveraging large-scale speech data to enhance prosodic modeling in speech synthesis, and introduces a model named SP2MC which achieves self-supervised prosody learning at phoneme-level with momentum contrast. This model incorporates dual convolutional encoders for speech and linear predictive coding (LPC) residual inputs to generate phoneme-level embeddings, which are masked and processed by a Transformer model to produce prosody representations. Two supervision modules are employed to generate phoneme-level supervision from speech waveforms and residuals. Momentum contrast is utilized to manage negative sample selection in contrastive learning. Finally, the SP2MC representations are integrated into a Fastspeech2-based acoustic model for speech synthesis. Experimental results indicate that the naturalness of speech synthesized by the proposed method is significantly better than that of baselines.

BibTeX
@inproceedings{icassp2025_selfsupervisedpr,
  title = {Self-supervised Prosody Learning at Phoneme-level with Momentum Contrast for Speech Synthesis},
  author = {Zhaoci Liu and Ya-Jun Hu and Liping Chen and Zhen-Hua Ling},
  booktitle = {ICASSP 2025},
  year = {2025}
}