2025
Towards Robust Autonomous Driving: Conditional Multimodal Large Language Models for Fine-Grained Perception
ICRA 2025
Multimodal large language models (MLLMs) have shown remarkable performance across various visual understanding tasks. However, most existing MLLMs still lack image detail perception, limiting their effectiveness in tasks that require detailed visual information. In this paper, we introduce Percept-D