2026
Towards Authentic Movie Dubbing with Retrieve-Augmented Director-Actor Interaction Learning
AAAI 2026technical
The automatic movie dubbing model generates vivid speech from given scripts, replicating a speaker
3 accepted papers
The automatic movie dubbing model generates vivid speech from given scripts, replicating a speaker
Conversational Speech Synthesis (CSS) aims to effectively take the multimodal dialogue history (MDH) to generate speech with appropriate conversational prosody for target utterance. The key challenge of CSS is to model the interaction between the MDH and the target utterance. Note that text and spee…
Conversational Speech Synthesis (CSS) aims to generate speech with natural prosody by understanding the multimodal dialogue history (MDH). The latest work predicts the accurate prosody expression of the target utterance by modeling the utterance-level interaction characteristics of MDH and the targe