CoSLR: Contrastive Chinese Sign Language Recognition with prior knowledge And Multi-Tasks Joint Learning
Tian Yang, Cong Shen, Tiantian Yuan
Abstract
Perceiving by computer vision, Sign Language Recognition (SLR) obtains the advantage of transforming the posture video into a sentence, compared with the methods of sensors to collect signals. However, learning representative features from a multimodal perspective is challenging. To this end, this study proposes a multi-task joint learning framework termed Contrastive Learning-based Sign Language Recognition Network (CoSLR) for Chinese sign language, which embeds text representation into the general video-based framework of SLR. In virtue of the profound ability of the pre-trained multimodal encoders, they are employed as processing modules to extract features from the original input video and text. Then, a contrastive learning between the video representation and corresponding token embedding is utilized to the feature extractor training. Finally, the linear combination of contrast and cross-entropy loss functions drives the end-to-end network to converge. Experiments show that the 1.27% WER of CoSLR has outperformed the state-of-the-art works in the comparison.
BibTeX
@inproceedings{icassp2024_coslrcontrastive,
title = {CoSLR: Contrastive Chinese Sign Language Recognition with prior knowledge And Multi-Tasks Joint Learning},
author = {Tian Yang and Cong Shen and Tiantian Yuan},
booktitle = {ICASSP 2024},
year = {2024}
}