Content-Based Objective Evaluation of Artificially Generated Sign Language Videos
Neha Tarigopula, Preyas Garg, Skanda Muralidhar, Sandrine Tornay, Dinesh Babu Jayagopi, Mathew Magimai-Doss
Abstract
Sign language is vital for communication within the deaf and hard-of-hearing community. Avatar-based methods and deep learning techniques like Generative Adversarial Networks have shown promise in generating sign language video content. One of the challenges in sign language generation is the evaluation of the generated video content. One possible solution is to subjectively evaluate using human raters. This is time-consuming and costly. The other possible solution is objective evaluation. In the literature, video quality metrics such as PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index) and skeleton-based measures such as MSE have been proposed. A limitation of these approaches is that they do not provide information about the generated video content. In this paper, we propose a novel phonology-based approach that evaluates the generated video along different channels, namely, hand movement and handshape, which convey the linguistic information in sign language. More precisely, in this approach an objective score is obtained by extracting sequences of hand movement sub-units and handshape sub-units class conditional probabilities (posterior features) from the source and generated videos and comparing them using dynamic time warping. Our experimental studies demonstrate that the proposed objective scoring method yields a better correlation to subjective human ratings than PSNR, SSIM, and MSE-based metrics.
BibTeX
@inproceedings{icassp2024_contentbasedobje,
title = {Content-Based Objective Evaluation of Artificially Generated Sign Language Videos},
author = {Neha Tarigopula and Preyas Garg and Skanda Muralidhar and Sandrine Tornay and Dinesh Babu Jayagopi and Mathew Magimai-Doss},
booktitle = {ICASSP 2024},
year = {2024}
}