ICRA 2026poster0 citations

Learning View-Invariant Sign Language Representations Via Dual-Stream Contrastive Learning

Yuting Peng, Yuecong Min, Xilin Chen

Abstract

Viewpoint shifts significantly change how gestures and facial expressions appear and frequently cause occlusions, posing a critical challenge for robust Sign Language Recognition (SLR). To address this challenge, we exploit the spatial flexibility and computational efficiency of skeleton data and propose ViSL, a dual-stream contrastive learning framework to learn underline{V}iew-underline{i}nvariant representations for underline{S}ign underline{L}anguage understanding. Specifically, the primary and lifting streams share a common visual feature extractor with different types of input: the primary stream (P-Stream) directly processes frontal-view skeleton data, and the lifting stream (L-Stream) synthesizes skeleton data from arbitrary viewpoints based on 3D estimations. We further propose a view-invariant contrastive loss to align representations across both viewpoints and streams. Experimental results on the challenging cross-view setting of MM-WLAuslan demonstrate that ViSL achieves substantial performance improvements, highlighting its potential for robust real-world SLR applications.

Gesture, Posture and Facial ExpressionsDeep Learning for Visual PerceptionRecognition
Learning View-Invariant Sign Language Representations Via Dual-Stream Contrastive Learning · ICRA 2026