Vision-Language Model Guided Semi-supervised Learning for No-Reference Video Quality Assessment
Shankhanil Mitra, Rajiv Soundararajan
Abstract
Perceptual assessment of user-generated content videos is an important problem that impacts viewing experience of millions of users. Current no-reference video quality assessment (NR-VQA) algorithms require a large amount of human annotated videos. In this work, we address this problem by specifically designing a dual-model based Semi-supervised Learning (SSL) method for NR-VQA. The first model is based on a popular vision-language model namely CLIP, where we adapt the visual encoder to capture high-level semantic quality information through quality-relevant text prompts. A second model learns complementary low-level spatio-temporal quality using a 3D vision transformer and video fragments. We enable intelligent knowledge transfer between the high-level vision-language and low-level vision-transformer model to pseudo-label the unlabelled videos. Our unified model outperforms existing state-of-the-art SSL methods for VQA across popular VQA databases including inter-database settings.
BibTeX
@inproceedings{icassp2025_visionlanguagemo,
title = {Vision-Language Model Guided Semi-supervised Learning for No-Reference Video Quality Assessment},
author = {Shankhanil Mitra and Rajiv Soundararajan},
booktitle = {ICASSP 2025},
year = {2025}
}