AAAI 2026technical0 citations

Scaling-up Perceptual Video Quality Assessment

Ziheng Jia, Zicheng Zhang, Xiaorong Zhu, Chunyi Li, Jinliang Han, Xiaohong Liu, Guangtao Zhai, Xiongkuo Min

Abstract

The data scaling law has significantly enhanced large multi-modal models (LMMs) performance across various downstream tasks. However, in the domain of perceptual video quality assessment (VQA), the potential of data scaling remains unprecedented due to the scarcity of labeled resources and the insufficient scale of datasets. To address this, we propose OmniVQA, a framework designed to efficiently build high-quality, machine-dominated synthetic multi-modal instruction databases (MIDBs) for VQA. We then scale up to create OmniVQA-Chat-400K, the largest dataset in the VQA field concurrently. Our focus is on the technical and aesthetic quality dimensions, with abundant in-context instruction data to provide fine-grained VQA knowledge. Additionally, we build the OmniVQA-MOS-20K dataset to enhance the model

BibTeX
@inproceedings{aaai2026_scalinguppercept,
  title = {Scaling-up Perceptual Video Quality Assessment},
  author = {Ziheng Jia and Zicheng Zhang and Xiaorong Zhu and Chunyi Li and Jinliang Han and Xiaohong Liu and Guangtao Zhai and Xiongkuo Min},
  booktitle = {AAAI 2026},
  year = {2026}
}
Scaling-up Perceptual Video Quality Assessment · AAAI 2026