ICASSP 2025accepted0 citations

FCConDubber: Fine And Coarse Grained Prosody Alignment For Expressive Video Dubbing via Contrastive Audio-Motion Pretraining

Qiulin Li, Zhichao Wu, Hanwei Li, Xin Dong, Qun Yang

Abstract

Automatic Video Dubbing (AVD) aims to synthesize speech that matches a character’s speaking style and emotion in silent video clips. However, existing approaches rely on attention mechanisms to learn cross-modal prosodic alignment implicitly, making it challenging to capture subtle prosodic variations and impacting the overall naturalness and expressiveness of the output. In this paper, we propose FCConDubber, a novel dubbing model that incorporates a contrastive speech-motion pre-training framework to learn fine-grained temporal prosodic alignment and coarse-grained global style information. Furthermore, we explore the relationship between facial features and audio features extracted from different layers of self-supervised speech representation models. Experiments on the MEAD dataset demonstrate that FCConDubber significantly outperforms baseline models in speech synthesis quality and prosody reconstruction. The synthesized samples are available at https://fccondubber.github.io/FCConDubber/.

BibTeX
@inproceedings{icassp2025_fccondubberfinea,
  title = {FCConDubber: Fine And Coarse Grained Prosody Alignment For Expressive Video Dubbing via Contrastive Audio-Motion Pretraining},
  author = {Qiulin Li and Zhichao Wu and Hanwei Li and Xin Dong and Qun Yang},
  booktitle = {ICASSP 2025},
  year = {2025}
}
FCConDubber: Fine And Coarse Grained Prosody Alignment For Expressive Video Dubbing via Contrastive Audio-Motion Pretraining · ICASSP 2025