ICRA 2024poster4 citations

Grounding Conversational Robots on Vision Through Dense Captioning and Large Language Models

Lucrezia Grassi, Zhouyang Hong, Carmine Tommaso Recchiuto, Antonio Sgorbissa

Abstract

This work explores a novel approach to empowering robots with visual perception capabilities using textual descriptions. Our approach involves the integration of GPT-4 with dense captioning, enabling robots to perceive and interpret the visual world through detailed text-based descriptions. To assess both user experience and the technical feasibility of this approach, experiments were conducted with human participants interacting with a Pepper robot equipped with visual capabilities. The results affirm the viability of the proposed approach, allowing to perform vision-based conversations effectively, despite processing time limitations.

BibTeX
@inproceedings{icra2024_groundingconvers,
  title = {Grounding Conversational Robots on Vision Through Dense Captioning and Large Language Models},
  author = {Lucrezia Grassi and Zhouyang Hong and Carmine Tommaso Recchiuto and Antonio Sgorbissa},
  booktitle = {ICRA 2024},
  year = {2024}
}
Grounding Conversational Robots on Vision Through Dense Captioning and Large Language Models · ICRA 2024