← Search

Songju Lei

2 accepted papers

2025

GLDiTalker: Speech-Driven 3D Facial Animation with Graph Latent Diffusion Transformer

IJCAI 2025

Speech-driven talking head generation is a critical yet challenging task with applications in augmented reality and virtual human modeling. While recent approaches using autoregressive and diffusion-based models have achieved notable progress, they often suffer from modality inconsistencies, particu

Cited by 0SourcePDFScholar
2024

Uni-Dubbing: Zero-Shot Speech Synthesis from Visual Articulation

ACL 2024long

In the field of speech synthesis, there is a growing emphasis on employing multimodal speech to enhance robustness. A key challenge in this area is the scarcity of datasets that pair audio with corresponding video. We employ a methodology that incorporates modality alignment during the pre-training…