IJCAI 20260 citations

G-VTM: A Multimodal Vision-Trajectory Model for Generalized Vehicle Trajectory Prediction

Xinyue Zhang, Letian Gong, Yan Lin, Jinjun Cheng, Junlin Zhang, Guanyu Yao, Shengnan Guo, Youfang Lin

Abstract

Generalized vehicle trajectory prediction across diverse junctions, including urban intersections and roundabouts, remains a fundamental task in Cooperative Vehicle–Infrastructure Systems (CVIS). This study faces two key challenges: (1) Generalize across junctions with heterogeneous map semantics and traffic behavioral patterns, where the former arises from differences in road topologies and traffic regulations, and the latter reflects diverse behavioral intentions of road users; (2) Scenario-adaptive interaction modeling, where single-modality trajectory learning captures local spatio-temporal correlation, but lacks map constraint and direction-aware interaction contexts. To overcome these challenges, we propose G-VTM, a generalized vision-trajectory model. G-VTM models fine-grained behavioral patterns and relative spatial interaction from trajectory modality. At the vision modality, G-VTM captures global map semantics while modeling scenario- and direction-aware interaction based on intuitive visual perception. Experiments on multiple real-world datasets collected by unmanned aerial vehicles (UAVs) demonstrate that our method achieves strong generalized performance under heterogeneous traffic conditions. The code is provided at https://github. com/zxyhaclyon/G-VTM.

Data Mining: Mining spatial and/or temporal data
BibTeX
@inproceedings{ijcai2026_gvtmamultimodalv,
  title = {G-VTM: A Multimodal Vision-Trajectory Model for Generalized Vehicle Trajectory Prediction},
  author = {Xinyue Zhang and Letian Gong and Yan Lin and Jinjun Cheng and Junlin Zhang and Guanyu Yao and Shengnan Guo and Youfang Lin and Shaojiang Wang and Huaiyu Wan},
  booktitle = {IJCAI 2026},
  year = {2026}
}
G-VTM: A Multimodal Vision-Trajectory Model for Generalized Vehicle Trajectory Prediction · IJCAI 2026