IROS 20250 citations

LGNav: Zero-Shot Object Navigation Driven by Language and Pointing Gesture Using Large Vision-Language Models

Weiyi Zhu, Juan Liu, Xinde Li, Zhiwei Lv, Zhehan Yang

Abstract

In human communication, referring to a specific object within an environment often involves the combination of a pointing gesture to indicate the object’s direction and linguistic descriptions specifying its name and attributes, thereby enabling precise object identification. Inspired by this natural multimodal interaction, we formalize the zero-shot object navigation driven by language and pointing gesture (LGZSON) task, which aims to more closely approximate real-world human-agent communication scenarios. To address this task, we propose LGNav, an open-set, training-free navigation framework. LGNav estimates the pointing gesture direction by extracting human body landmarks and integrates this directional information with depth images to initialize a versatile candidate position map (VCPM). The framework further employs open-vocabulary object detection to identify all potential candidate objects in the environment, projecting them onto the VCPM. Guided by a motion policy derived from the VCPM, LGNav continuously explores the unknown environment, sequentially visits candidate objects, and utilizes a large vision-language model (LVLM) to verify whether each candidate object satisfies the given navigation instruction. Extensive experimental results validate the effectiveness of LGNav, demonstrating its strong performance in the LG-ZSON task. Furthermore, even in the absence of pointing gestures, LGNav achieves competitive results on standard object navigation benchmarks, including the Gibson and HM3D datasets, outperforming a range of strong baseline methods.

BibTeX
@inproceedings{iros2025_lgnavzeroshotobj,
  title = {LGNav: Zero-Shot Object Navigation Driven by Language and Pointing Gesture Using Large Vision-Language Models},
  author = {Weiyi Zhu and Juan Liu and Xinde Li and Zhiwei Lv and Zhehan Yang},
  booktitle = {IROS 2025},
  year = {2025}
}
LGNav: Zero-Shot Object Navigation Driven by Language and Pointing Gesture Using Large Vision-Language Models · IROS 2025