FCConDubber: Fine And Coarse Grained Prosody Alignment For Expressive Video Dubbing via Contrastive Audio-Motion Pretraining
Qiulin Li, Zhichao Wu, Hanwei Li, Xin Dong, Qun Yang
Abstract
Automatic Video Dubbing (AVD) aims to synthesize speech that matches a character’s speaking style and emotion in silent video clips. However, existing approaches rely on attention mechanisms to learn cross-modal prosodic alignment implicitly, making it challenging to capture subtle prosodic variations and impacting the overall naturalness and expressiveness of the output. In this paper, we propose FCConDubber, a novel dubbing model that incorporates a contrastive speech-motion pre-training framework to learn fine-grained temporal prosodic alignment and coarse-grained global style information. Furthermore, we explore the relationship between facial features and audio features extracted from different layers of self-supervised speech representation models. Experiments on the MEAD dataset demonstrate that FCConDubber significantly outperforms baseline models in speech synthesis quality and prosody reconstruction. The synthesized samples are available at https://fccondubber.github.io/FCConDubber/.
BibTeX
@inproceedings{icassp2025_fccondubberfinea,
title = {FCConDubber: Fine And Coarse Grained Prosody Alignment For Expressive Video Dubbing via Contrastive Audio-Motion Pretraining},
author = {Qiulin Li and Zhichao Wu and Hanwei Li and Xin Dong and Qun Yang},
booktitle = {ICASSP 2025},
year = {2025}
}