2022
Interactive Multi-Level Prosody Control for Expressive Speech Synthesis
ICASSP 2022accepted
Recent neural-based text-to-speech (TTS) models are able to produce highly natural speech. To synthesize expressive speech, the prosody of the speech has to be modeled, and predicted/controlled during synthesis. However, intuitive control over prosody remains elusive. Some techniques only allow cont…