ICRA 2026poster0 citations

V2V-GoT: Vehicle-To-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models and Graph-Of-Thoughts

Hsu-kuang Chiu, Ryo Hachiuma, Chien-Yi Wang, Yu-Chiang Frank Wang, Min-Hung Chen, Stephen F. Smith

Abstract

Current state-of-the-art autonomous vehicles could face safety critical situations when their local sensors are occluded by large objects on the road nearby. Vehicle-to-vehicle (V2V) cooperative autonomous driving is proposed to address this problem. More recent work further adopts a new approach that applies Multimodal Large Language Models (MLLMs) for cooperative autonomous driving due to its potential multimodal understanding and reasoning abilities. However, graph-of-thoughts reasoning frameworks have not been considered for prior research on V2V cooperative autonomous driving. In this paper, we propose a novel graph-of-thoughts framework specifically designed for MLLM-based cooperative autonomous driving. Our graph-of-thoughts includes our proposed novel ideas of occlusion-aware perception and planning-aware prediction. We curate the V2V-GoT-QA dataset and develop the V2V-GoT model for training and testing the cooperative driving graph-of-thoughts. Our experimental results show that our proposed method outperforms other baselines in cooperative perception, prediction, and planning tasks. Our code and dataset are released to facilitate open-source research at https://eddyhkchiu.github.io/v2vgot.github.io/.

Computer Vision for TransportationIntelligent Transportation SystemsDeep Learning for Visual Perception