ACL 2025finding0 citations

TC-Bench: Benchmarking Temporal Compositionality in Conditional Video Generation

Weixi Feng, Jiachen Li, Michael Saxon, Tsu-Jui Fu, Wenhu Chen, William Yang Wang

Abstract

Video generation has many unique challenges beyond those of image generation. The temporal dimension introduces extensive possible variations across frames, over which consistency and continuity may be violated. In this work, we evaluate the emergence of new concepts and relation transitions as time progresses in a video, which we refer to as Temporal Compositionality. We propose TC-Bench, a benchmark of meticulously crafted text prompts, ground truth videos, and new evaluation metrics. The prompts articulate the initial and final states of scenes, effectively reducing ambiguities for frame development. In addition, by collecting corresponding ground-truth videos, the benchmark can be used for text-to-video and image-to-video generation. We develop new metrics to measure the completeness of component transitions, which demonstrate significantly higher correlations with human judgments than existing metrics. Our experiments reveal that contemporary video generators are still weak in prompt understanding and achieve less than 20% of the compositional changes, highlighting enormous improvement space. Our analysis indicates that current video generation models struggle to interpret descriptions of compositional changes and synthesize various components across different time steps.

BibTeX
@inproceedings{feng-etal-2025-tc,
    title = "{TC}-Bench: Benchmarking Temporal Compositionality in Conditional Video Generation",
    author = "Feng, Weixi  and
      Li, Jiachen  and
      Saxon, Michael  and
      Fu, Tsu-Jui  and
      Chen, Wenhu  and
      Wang, William Yang",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2025",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.findings-acl.241/",
    doi = "10.18653/v1/2025.findings-acl.241",
    pages = "4638--4662",
    ISBN = "979-8-89176-256-5"
}
TC-Bench: Benchmarking Temporal Compositionality in Conditional Video Generation · ACL 2025