TagGuideBot: Enhancing Robot Intelligence with Object Tags and VLMs
Jiayi Chen, Ying He, F. Richard Yu
Abstract
This research aims to enhance the interaction between humans and robots, especially in environments with multiple similar objects or semantic ambiguities. Traditional command-based interactions typically require users to provide precise descriptions, which often poses a significant challenge. To address this issue, we propose a framework named Tag-GuideBot, which leverages Visual Language Models (VLMs) and utilizes object markers to help locate and identify objects in the environment. By integrating positional point prompts of the target objects with robot motion planning models, we aim to achieve a more accurate understanding and execution of complex commands, thus improving the efficiency and naturalness of interactions. Experimental results demonstrate that TagGuideBot effectively addresses the challenges posed by complex commands and environmental complexities, achieving an accuracy of 66.3% on user instructions extended beyond the training set, providing solid support for further optimization of human-robot interaction.
BibTeX
@inproceedings{iros2025_tagguidebotenhan,
title = {TagGuideBot: Enhancing Robot Intelligence with Object Tags and VLMs},
author = {Jiayi Chen and Ying He and F. Richard Yu},
booktitle = {IROS 2025},
year = {2025}
}