RA-L 20261 citations

Curiosity-Driven Zero-Shot Object Navigation With Vision-Language Models

Zhengyi Gu, Yuheng Zhou, Yuquan Xue, Hanzhang Wang

Abstract

Zero-shot object navigation (ZSON) in unseen environments poses a significant challenge due to the absence of object-specific priors and the need for efficient exploration. Existing approaches often struggle with ineffective search strategies and repeated visits to irrelevant areas. In this paper, we introduce a curiosity-driven framework that leverages the commonsense reasoning capabilities of vision-language models (VLMs) to guide exploration. At each step, the agent estimates the semantic plausibility of regions based on language-conditioned visual cues, constructing a dynamic value map that promotes informative regions and suppresses redundancy. The core contribution of this work is integrating VLM-based scene understanding into the curiosity mechanism, enabling the agent to make human-like judgments about environmental relevance during navigation. Extensive experiments on the HM3D benchmark show that our method achieves a 12.1% absolute improvement in Success Rate (SR) over strong baselines (from 56.5% to 68.6%). Qualitative analysis further confirms that the proposed strategy leads to more efficient and goal-directed exploration.

BibTeX
@inproceedings{ral2026_curiositydrivenz,
  title = {Curiosity-Driven Zero-Shot Object Navigation With Vision-Language Models},
  author = {Zhengyi Gu and Yuheng Zhou and Yuquan Xue and Hanzhang Wang},
  booktitle = {RA-L 2026},
  year = {2026}
}