Curiosity-Driven Zero-Shot Object Navigation With Vision-Language Models
Zhengyi Gu, Yuheng Zhou, Yuquan Xue, Hanzhang Wang
Abstract
Zero-shot object navigation (ZSON) in unseen environments poses a significant challenge due to the absence of object-specific priors and the need for efficient exploration. Existing approaches often struggle with ineffective search strategies and repeated visits to irrelevant areas. In this paper, we introduce a curiosity-driven framework that leverages the commonsense reasoning capabilities of vision-language models (VLMs) to guide exploration. At each step, the agent estimates the semantic plausibility of regions based on language-conditioned visual cues, constructing a dynamic value map that promotes informative regions and suppresses redundancy. The core contribution of this work is integrating VLM-based scene understanding into the curiosity mechanism, enabling the agent to make human-like judgments about environmental relevance during navigation. Extensive experiments on the HM3D benchmark show that our method achieves a 12.1% absolute improvement in Success Rate (SR) over strong baselines (from 56.5% to 68.6%). Qualitative analysis further confirms that the proposed strategy leads to more efficient and goal-directed exploration.
BibTeX
@inproceedings{ral2026_curiositydrivenz,
title = {Curiosity-Driven Zero-Shot Object Navigation With Vision-Language Models},
author = {Zhengyi Gu and Yuheng Zhou and Yuquan Xue and Hanzhang Wang},
booktitle = {RA-L 2026},
year = {2026}
}