AAAI 2026technical0 citations

VMChill: A Dataset for Fine-Grained Visual-Musical Synergy

Xiaowei Chi, Zeyue Tian, Jialiang Chen, Wei Xue

Abstract

Massive multi-modality datasets are fundamental to the success of large video-language models. However, existing datasets often focus on providing textual descriptions for visual content, treating audio, particularly music, as weakly related information. This overlooks the inherent semantic correlation between visual narratives and musical scores, limiting the development of models for fine-grained cross-modal understanding and generation. To address this gap, we introduce VMChill, a large-scale, fine-grained multimodal video dataset. We leverage trailers as our data source, as they are professionally edited to create a strong synergy between visual pacing, scene transitions, and background music for narrative and emotional impact. Our dataset comprises over 20 million video clips derived from more than 27.1k hours of high-resolution trailer videos. To annotate this data, we propose a systematic multimodal captioning framework. This framework first employs specialized unimodal models to extract descriptive features from multiple perspectives, including visual content, motion dynamics, and musical attributes (e.g., genre, instruments, mood). Subsequently, a large language model (LLM) is utilized to adaptively fuse these diverse descriptions into a single, coherent, and rich multimodal caption. This process yields VMChill-2M, a high-quality subset of 2 million clips with detailed multimodal annotations, and VMChill-Test, a manually refined test set for evaluation. We conduct extensive experiments on downstream tasks, including video understanding and generation, to establish benchmarks and demonstrate the dataset

BibTeX
@inproceedings{aaai2026_vmchilladatasetf,
  title = {VMChill: A Dataset for Fine-Grained Visual-Musical Synergy},
  author = {Xiaowei Chi and Zeyue Tian and Jialiang Chen and Wei Xue},
  booktitle = {AAAI 2026},
  year = {2026}
}
VMChill: A Dataset for Fine-Grained Visual-Musical Synergy · AAAI 2026