CVPR 2024poster35 citations

On the Content Bias in Frechet Video Distance

Songwei Ge, Aniruddha Mahapatra, Gaurav Parmar, Jun-Yan Zhu, Jia-Bin Huang

Abstract

Frechet Video Distance (FVD) a prominent metric for evaluating video generation models is known to conflict with human perception occasionally. In this paper we aim to explore the extent of FVD's bias toward frame quality over temporal realism and identify its sources. We first quantify the FVD's sensitivity to the temporal axis by decoupling the frame and motion quality and find that the FVD only increases slightly with larger temporal corruption. We then analyze the generated videos and show that via careful sampling from a large set of generated videos that do not contain motions one can drastically decrease FVD without improving the temporal quality. Both studies suggest FVD's basis towards the quality of individual frames. We show that FVD with features extracted from the recent large-scale self-supervised video models is less biased toward image quality. Finally we revisit a few real-world examples to validate our hypothesis.

BibTeX
@inproceedings{cvpr2024_onthecontentbias,
  title = {On the Content Bias in Frechet Video Distance},
  author = {Songwei Ge and Aniruddha Mahapatra and Gaurav Parmar and Jun-Yan Zhu and Jia-Bin Huang},
  booktitle = {CVPR 2024},
  year = {2024}
}
On the Content Bias in Frechet Video Distance · CVPR 2024