CI-VID: A Coherent Interleaved Text-Video Dataset
Yiming Ju, Jijin Hu, Zhengxiong Luo, Haoge Deng, hanyu Zhao, Li Du, Wenbo Xiao, Chengwei Wu
Abstract
Text-to-video (T2V) generation has recently attracted considerable attention, resulting in the development of numerous high-quality datasets that have propelled progress in this area. However, existing public datasets are primarily composed of isolated text-video (T-V) pairs and thus fail to model inter-clip relationships. To address this limitation, we introduce CI-VID, a dataset that moves beyond isolated T2V generation toward text-and-video-to-video (T&V2V) generation. CI-VID contains over 340,000 samples, each comprising a semantically coherent video sequence with interleaved text captions that capture both clip-level content and inter-clip relationships. To validate its effectiveness, we design a comprehensive, multi-dimensional benchmark incorporating human evaluation, VLM-based assessment, and similarity-based metrics. Experimental results demonstrate that models trained on CI-VID significantly improve both accuracy and content consistency in multi-clip video generation. This enables the creation of story-driven content with smooth transitions and strong semantic coherence.
BibTeX
@inproceedings{cvpr2026_cividacoherentin,
title = {CI-VID: A Coherent Interleaved Text-Video Dataset},
author = {Yiming Ju and Jijin Hu and Zhengxiong Luo and Haoge Deng and hanyu Zhao and Li Du and Wenbo Xiao and Chengwei Wu and Donglin Hao and Xinlong Wang and Tengfei Pan},
booktitle = {CVPR 2026},
year = {2026}
}