Multi-Modal Rhythmic Generative Model for Chinese Cued Speech Gestures Generation
Li Liu, Wentao Lei, Wenwu Wang
Abstract
Cued speech (CS) is a novel visual coding system, which combines lip reading with several specific hand codings to help hearing-impaired people to communicate effectively. This work focuses on the audio/text-driven CS gestures (i.e., continuous lip and hand gestures movements) generation. Previous work used template-based statistical methods for the French CS generation. However, these methods are fragile since they need careful hand-crafted pre-processing to fit models, resulting in poor robustness. Furthermore, the natural rhythm in generated CS gesture sequences, which is essential for a coding system of spoken languages, was overlooked in prior studies. To solve the above-mentioned problems, we innovatively propose a two-branched rhythmic CS gesture generation framework, which contains a multi-modal adversarial semantic generator (MASG) to generate accurate multi-modal CS gestures (i.e., lip, hand shape and hand position movements), and an audio-driven rhythm generator (ARG) to extract the rhythm information. Moreover, we design a new Gesture Audio Difference (GAD) metric to evaluate the rhythm coherence considering the issue of asynchrony between CS hand gestures and lip movements. Extensive experimental results are presented on two datasets of two tasks (a CS dataset named MCCS-2024 and a co-speech TED dataset) with comprehensive ablation analysis and user study, demonstrating the effectiveness of our method. The code and dataset with multi-modal annotations were made public at https://mccs-2024.github.io/.
BibTeX
@inproceedings{icassp2025_multimodalrhythm,
title = {Multi-Modal Rhythmic Generative Model for Chinese Cued Speech Gestures Generation},
author = {Li Liu and Wentao Lei and Wenwu Wang},
booktitle = {ICASSP 2025},
year = {2025}
}