Enhancing Prosody Transfer in Speech Synthesis by Using Prosodically-Aligned References
Abstract
Prosody transfer in speech synthesis aims to produce natural and expressive speech by replicating the prosody of reference speech. Traditional methods, which rely on same-text, samespeaker references during training, often fail to generalize effectively when applied to different-speaker or different-text scenarios during inference, resulting in degraded quality and speaker leakage. To address these limitations, we propose a novel method that leverages content- and speaker-independent references during training, which are prosodically-aligned—meaning that they are closely matched in rhythm, intonation, and stress patterns with the target speech. Specifically, this approach employs non-target references in training to closely mirror test-time conditions. To ensure effective prosody transfer, unit selection is utilized to choose and concatenate segments that closely match the prosody of the target utterance. Additionally, speaker-specific features are carefully normalized during the target cost computation in the unit selection process to enhance the preservation of the target speaker’s identity. Evaluation results demonstrate that our method achieves more consistent performance between training and inference, better preserves the target speaker identity, and generates prosody that is comparable to models trained with ground truth references.
BibTeX
@inproceedings{icassp2025_enhancingprosody,
title = {Enhancing Prosody Transfer in Speech Synthesis by Using Prosodically-Aligned References},
author = {Lin Liu},
booktitle = {ICASSP 2025},
year = {2025}
}