ICML 2025poster0 citations

TUMTraf VideoQA: Dataset and Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes

Xingcheng Zhou, Konstantinos Larintzakis, Hao Guo, Walter Zimmer, Mingyu Liu, Hu Cao, Jiajie Zhang, Venkatnarayanan Lakshminarasimhan

Abstract

We present TUMTraf VideoQA, a novel dataset and benchmark designed for spatio-temporal video understanding in complex roadside traffic scenarios. The dataset comprises 1,000 videos, featuring 85,000 multiple-choice QA pairs, 2,300 object captioning, and 5,700 object grounding annotations, encompassing diverse real-world conditions such as adverse weather and traffic anomalies. By incorporating tuple-based spatio-temporal object expressions, TUMTraf VideoQA unifies three essential tasks—multiple-choice video question answering, referred object captioning, and spatio-temporal object grounding—within a cohesive evaluation framework. We further introduce the TraffiX-Qwen baseline model, enhanced with visual token sampling strategies, providing valuable insights into the challenges of fine-grained spatio-temporal reasoning. Extensive experiments demonstrate the dataset’s complexity, highlight the limitations of existing models, and position TUMTraf VideoQA as a robust foundation for advancing research in intelligent transportation systems. The dataset and benchmark are publicly available to facilitate further exploration.

Traffic Scene UnderstandingVideo UnderstandingVision Language ModelSpatial Temporal ReasoningAutonomous Driving
BibTeX
@inproceedings{
zhou2025tumtraf,
title={{TUMT}raf Video{QA}: Dataset and Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes},
author={Xingcheng Zhou and Konstantinos Larintzakis and Hao Guo and Walter Zimmer and Mingyu Liu and Hu Cao and Jiajie Zhang and Venkatnarayanan Lakshminarasimhan and Leah Strand and Alois Knoll},
booktitle={Forty-second International Conference on Machine Learning},
year={2025},
url={https://openreview.net/forum?id=Yfoi5O68rf}
}