← Search

Qiulin Li

2 accepted papers

2025

FCConDubber: Fine And Coarse Grained Prosody Alignment For Expressive Video Dubbing via Contrastive Audio-Motion Pretraining

ICASSP 2025accepted

Automatic Video Dubbing (AVD) aims to synthesize speech that matches a character’s speaking style and emotion in silent video clips. However, existing approaches rely on attention mechanisms to learn cross-modal prosodic alignment implicitly, making it challenging to capture subtle prosodic variatio…

Cited by 0SourceScholar
2024

DCTTS: Discrete Diffusion Model with Contrastive Learning for Text-to-Speech Generation

ICASSP 2024accepted

In the Text-to-speech(TTS) task, the latent diffusion model has excellent fidelity and generalization, but its expensive resource consumption and slow inference speed have always been a challenging. To address this issue, this paper proposes the Discrete Diffusion Model with Contrastive Learning for…

Cited by 0SourceScholar