VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction
Hao Wang, Eiki Murata, Lingfang Zhang, Ayako Sato, So Fukuda, Ziqi Yin, Wentao Hu, Keisuke Nakao
Abstract
Recent advances in multimodal large language models (MLLMs) have significantly enhanced video understanding capabilities, opening new possibilities for practical applications. Yet current video benchmarks focus largely on indoor scenes or short-range outdoor activities, leaving the challenges associated with long-distance travel largely unexplored. Mastering extended geospatial-temporal trajectories is critical for next-generation MLLMs, underpinning real-world tasks such as embodied-AI planning and navigation. To bridge this gap, we present VIR-Bench, a novel benchmark consisting of 200 travel videos that frames itinerary reconstruction as a challenging task designed to evaluate and push forward MLLMs
BibTeX
@inproceedings{aaai2026_virbenchevaluati,
title = {VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction},
author = {Hao Wang and Eiki Murata and Lingfang Zhang and Ayako Sato and So Fukuda and Ziqi Yin and Wentao Hu and Keisuke Nakao and Yusuke Nakamura and Sebastian Zwirner and Yi-Chia Chen and Hiroyuki Otomo and Hiroki Ouchi and Daisuke Kawahara},
booktitle = {AAAI 2026},
year = {2026}
}