Inter- and Intra-Sentence Cuer-Invariant Representation Learning for Generalizable Cued Speech Recognition
Abstract
Cued Speech (CS) is a visual coding system that combines lip movements and hand gestures to represent spoken languages for hearing-impaired people. Automatic Cued Speech Recognition (ACSR) is an emerging research topic, but the cuer (i.e., people who perform CS) generalization problem of ACSR remains unexplored. Moreover, the asynchrony between lip and hand modalities further aggravates the challenges associated with the generalization problem. Therefore, we propose a novel multi-modal, cuer-invariant representation learning framework that facilitates generalizable ACSR through contrastive learning, enabling our model to recognize unseen cuers. The proposed approach comprises three key components: 1) an inter-sentence contrastive learning module that learns cuer-invariant representations for lip and hand modalities to solve the cuer generalization problem in ACSR; and 2) a multi-level intra-sentence cross-attention module that synchronizes the lip movements and hand gestures to address the modality asynchrony issue in ACSR; 3) an easy-to-hard progressive learning strategy to stabilize the learning process and prevent performance degradation on hard examples. Extensive experiments on two available CS datasets show that our method outperforms previous works by a large margin. We also conduct ablation studies and visualization to demonstrate the effectiveness of the proposed method.
BibTeX
@inproceedings{icassp2025_interandintrasen,
title = {Inter- and Intra-Sentence Cuer-Invariant Representation Learning for Generalizable Cued Speech Recognition},
author = {Tianxin Xie and Li Liu},
booktitle = {ICASSP 2025},
year = {2025}
}