Synthesizing Efficient Trajectory-Controllable Co-Speech Gesture with Latent Consistency Model
Wenlong Wang, Dahua Gao, Yuxi Hu
Abstract
Diffusion models excel at audio-driven gesture generation. However, the reverse denoising process is computationally intensive and time-redundant. Additionally, existing methods predominantly focus on the quality and diversity of upper-body gestures, without considering the movement trajectories of speaker. To address these issues, we introduce, for the first time, a Latent Consistency Model for co-speech gesture generation (GestureLCM), which allows for large-scale skip sampling while ensuring consistency between adjacent perturbations in the latent probability flow space. Moreover, we integrate ControlNet into the latent space of GestureLCM to explicitly control the trajectory generation using spatial signals. By employing this approach, our model can generate content-related movements. Extensive experiments show the efficiency and controllability of our method for audio-driven co-speech gesture generation, achieving 50× acceleration in inference speed.
BibTeX
@inproceedings{icassp2025_synthesizingeffi,
title = {Synthesizing Efficient Trajectory-Controllable Co-Speech Gesture with Latent Consistency Model},
author = {Wenlong Wang and Dahua Gao and Yuxi Hu},
booktitle = {ICASSP 2025},
year = {2025}
}