Vision-Language Guided Adaptive Robot Action Planning: Responding to Intermediate Results and Implicit Human Intentions
Weihao Cai, Yoshiki Mori, Nobutaka Shimada
Abstract
Recent advances in research have demonstrated that Vision-Language Models (VLMs) are a promising technology for robot task planning. This paper presents a novel approach that leverages visual prompts and VLMs to generate feasible robot action sequences for achieving shared tasks through human-robot collaboration while simultaneously estimating human intentions. Our method enhances VLMs’ understanding of the environment by utilizing annotations (bounding boxes and labels) and dynamically infers human intentions based on changing environmental conditions to generate optimal robot action sequences to achieve common goals. Additionally, the system incorporates a mechanism to regenerate new sequences through VLM analysis when action failures or external interference occur. Furthermore, by designing prompts as versatile modules for diverse tasks, our proposed technology offers a new approach to robot action planning that excels in both efficiency and adaptability.
BibTeX
@inproceedings{iros2025_visionlanguagegu,
title = {Vision-Language Guided Adaptive Robot Action Planning: Responding to Intermediate Results and Implicit Human Intentions},
author = {Weihao Cai and Yoshiki Mori and Nobutaka Shimada},
booktitle = {IROS 2025},
year = {2025}
}