CVPR 20260 citations

Ref4D-VideoBench: Four-Dimensional Reference-Based Evaluation of Text-to-Video Generative Models

Jiajia Wei, Yujia He, Yuhan Hou, Hang Qi, Sihua Wang, Jincheng Shi, Kwok Fung Li, Zibin Zheng

Abstract

Most existing evaluations of generated videos adopt a no-reference paradigm. Although recent benchmarks cover multiple dimensions and show moderate correlation with human preferences, relying solely on textual prompts weakens real-world constraints and makes it difficult to produce accountable and interpretable judgments on instance-level issues such as target behavior deviation, temporal inconsistency, and commonsense violations. In scenarios with explicit expectations, such as controlled generation, reference videos naturally provide rich, unambiguous spatio-temporal evidence, enabling stricter and more trustworthy assessment. Motivated by this, we propose Ref4D-VideoBench, a reference-based, fine-grained, multi-dimensional benchmark for generated video evaluation. Ref4D-VideoBench contains 600 high-quality reference videos with tightly evidence-bounded prompts, and introduces a 12-metric structured evaluation suite along four key dimensions: basic semantic alignment, motion consistency, event temporal consistency, and world knowledge consistency. Experiments on eight text-to-video models show that our method achieves stronger agreement with human judgments than representative no-reference frameworks. Our code is available at https://github.com/TAILab-W/Ref4D-VideoBench.

BibTeX
@inproceedings{cvpr2026_ref4dvideobenchf,
  title = {Ref4D-VideoBench: Four-Dimensional Reference-Based Evaluation of Text-to-Video Generative Models},
  author = {Jiajia Wei and Yujia He and Yuhan Hou and Hang Qi and Sihua Wang and Jincheng Shi and Kwok Fung Li and Zibin Zheng and Weibin Wu},
  booktitle = {CVPR 2026},
  year = {2026}
}
Ref4D-VideoBench: Four-Dimensional Reference-Based Evaluation of Text-to-Video Generative Models · CVPR 2026