Hierarchical Spatial-Temporal Enhancement Network For Continuous Sign Language Recognition
In continuous sign language recognition (CSLR), 2D-CNN-based extractors are often insufficiently trained for spatial capture and struggle with temporal modeling. This leads to incomplete spatial discrimination, hindering the understanding actions across frames. To address these limitations, we propo…