2025
Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models
IROS 2025
In this paper, we present a hierarchical question-answering (QA) approach for scene understanding in autonomous vehicles, balancing cost-efficiency with detailed visual interpretation. The method fine-tunes a compact vision-language model (VLM) on a custom dataset specific to the geographical area i