MTRDrive: Memory-Tool Synergistic Reasoning for Robust Autonomous Driving in Corner Cases
Ziang Luo, Kangan Qian, Jiahua Wang, Jinyu Miao, Zheng Fu, Yuechen Luo, Yunlong Wang, Sicong Jiang
Abstract
Vision-Language Models (VLMs) have demonstrated significant potential for end-to-end autonomous driving, yet a substantial gap remains between their current capabilities and the reliability necessary for real-world deployment. A critical challenge is their fragility, characterized by hallucinations and poor generalization in out-of-distribution (OOD) scenarios. To bridge this gap, we introduce MTRDrive, a novel framework that integrates procedural driving experiences with a dynamic toolkit to enhance generalization and proactive decision-making. MTRDrive addresses these limitations through a closed-loop system that combines a memory-based experience retrieval mechanism with dynamic toolkits. This synergy enables the model to interact more effectively with its environment, improving both reasoning and decision-making capabilities with the help of our memory-tool synergistic reasoning. Additionally, we introduce a new benchmark based on complex Roadwork construction scenarios to rigorously evaluate zero-shot generalization. Extensive experiments demonstrate the effectiveness of our approach. On the public NavSim benchmark, MTRDrive achieves state-of-the-art performance with a driving metric score of 79.8% and a planning accuracy of 82.6%. To rigorously test generalization, we evaluate our model in a zero-shot setting on our new Roadwork-VLM benchmark. In this challenging out-of-distribution test, it attains a driving metric score of 80.2% and a planning accuracy of 33.5%, showcasing its strong ability to reason robustly in unseen scenarios. These results highlight the potential of MTRDrive to advance the field of autonomous driving towards safer and more reliable systems.