A Gloss-Driven Indian Sign Language Production System Using Learned Pose Representations
Suvajit Patra, Arkadip Maitra, Swami Punyeshwarananda, Soumitra Samanta
Abstract
Sign Language Production (SLP) system translates spoken or written language into sign language, enabling accessible communication between the deaf/hard-of-hearing and the hearing population. Being one of the most widely used sign languages globally, Indian Sign Language (ISL) is a very low-resource language and lacks such SLP systems. This paper presents a scalable and modular SLP framework based on Sign-Pose-VQ-VAE model, designed for low-resource settings. The model learns discrete pose representations (codes) by disentangling body, left-hand, and right-hand keypoints, enabling efficient pose modeling and co-articulated sign generation. The proposed system is evaluated using a Hindi movie subtitle corpus coupled with an off-the-shelf back-translation model and achieves a gloss BLEU-4 score of 47.20. The system generated signs are evaluated by certified ISL interpreters with an average rating of 4.33/5, and a BERT precision of 0.7683 on glosses. In addition, the proposed system achieves state-of-the-art performance among keypoint-based methods on the PHOENIX14T benchmark, with a BLEU-4 score of 10.03.
BibTeX
@inproceedings{ijcai2026_aglossdrivenindi,
title = {A Gloss-Driven Indian Sign Language Production System Using Learned Pose Representations},
author = {Suvajit Patra and Arkadip Maitra and Swami Punyeshwarananda and Soumitra Samanta},
booktitle = {IJCAI 2026},
year = {2026}
}