IROS 20250 citations

Dual-Level Open-Vocabulary 3D Scene Representation for Instance-Aware Robot Navigation

Tianlu Zheng, Kaicheng Yang, Yilong Dou, Ziyong Feng, Qichuan Ding

Abstract

Advanced scene understanding is crucial for robots to navigate robustly in complex 3D environments. Recent works utilize large Vision-Language Models (VLMs) to embed semantic information into reconstructed maps, thereby creating open-vocabulary scene representations for instance-aware robot navigation. However, existing methods primarily generate point-wise feature vectors for maps, which inadequately capture the intricate scene contents necessary for navigation tasks, including holistic and relational object information. To address this limitation, we propose a novel Dual-Level Open-Vocabulary 3D (DLOV-3D) scene representation framework to improve robot navigation performance. Our framework integrates both pixel-level and image-level features into spatial scene representations, facilitating a more comprehensive understanding of the scene. By incorporating an adaptive revalidation mechanism, DLOV-3D achieves precise instance-aware navigation based on free-form queries that describe object properties such as color, shape, and relational references. Notably, when combined with Large Language Models (LLMs), DLOV-3D supports long-sequence multi-instance robot navigation guided by natural language instructions. Extensive experimental results demonstrate that DLOV-3D achieves new state-of-the-art performance in instance-aware robot navigation.

BibTeX
@inproceedings{iros2025_duallevelopenvoc,
  title = {Dual-Level Open-Vocabulary 3D Scene Representation for Instance-Aware Robot Navigation},
  author = {Tianlu Zheng and Kaicheng Yang and Yilong Dou and Ziyong Feng and Qichuan Ding},
  booktitle = {IROS 2025},
  year = {2025}
}