ICRA 2026poster0 citations

OccLLaMA: A Unified Occupancy-Language-Action World Model for Enhancing Motion Planning Via Multi-Task Learning

Julong Wei, Shanshuai Yuan, Pengfei Li, Xinyi Quan, Lei Tai, Jieru Zhao, Zhongxue Gan, Wenchao Ding

Abstract

Scene understanding via multi-modal large language models and scene forecasting with world models have advanced the development of autonomous driving. The former maps visual inputs to driving-specific outputs, neglecting spatial reasoning and world dynamics. The latter captures world dynamics, lacking comprehensive scene understanding. In contrast, human divers seamlessly integrate understanding, forecasting, and decision-making through multi-modal representations. To this end, we propose OccLLaMA, a unified occupancy-language-action world model to enhance motion planning via multi-task learning. It uses semantic occupancy as a unified 3D visual representation, effectively integrating spatial scene understanding and forecasting. Specifically, we first introduce a tailored scene tokenizer that auto-encodes semantic occupancy into latent tokens for invertible compression. Furthermore, we enhance LLaMA to enable joint learning across both understanding and generation tasks within a unified auto-regressive framework, incorporating multi-task pretraining and motion-planning–oriented fine-tuning. Extensive experiments demonstrate that OccLLaMA not only achieves competitive performance on scene understanding and occupancy forecasting, but also enhances motion planning by integrating multi-task inference, showcasing its effectiveness and potential as a foundation model for autonomous driving. Project page: href{https://vilonge.github.io/OccLLaMA_Page/}{OccLLaMA}

Computer Vision for AutomationDeep Learning for Visual PerceptionAutonomous Vehicle Navigation
OccLLaMA: A Unified Occupancy-Language-Action World Model for Enhancing Motion Planning Via Multi-Task Learning · ICRA 2026