2026
NO VERIFIABLE REWARD FOR PROSODY: TOWARD PREFERENCE-GUIDED PROSODY LEARNING IN TTS
ICASSP 2026poster
Recent work reports gains in neural text-to-speech (TTS) with Group Relative Policy Optimization (GRPO). However, in the absence of a verifiable reward for \textit{prosody}, GRPO trained on transcription-oriented signals (CER/NLL) lowers error rates yet collapses prosody into monotone, unnatural spe…