2025
Democratizing High-Fidelity Co-Speech Gesture Video Generation
ICCV 2025poster
Co-speech gesture video generation aims to synthesize realistic, audio-aligned videos of speakers, complete with synchronized facial expressions and body gestures. This task presents challenges due to the significant one-to-many mapping between audio and visual content, further complicated by the sc…