CoRL 2024poster12 citations

A3VLM: Actionable Articulation-Aware Vision Language Model

Siyuan Huang, Haonan Chang, Yuhan Liu, Yimeng Zhu, Hao Dong, Abdeslam Boularias, Peng Gao, Hongsheng Li

Abstract

Vision Language Models (VLMs) for robotics have received significant attention in recent years. As a VLM can understand robot observations and perform complex visual reasoning, it is regarded as a potential universal solution for general robotics challenges such as manipulation and navigation. However, previous robotics VLMs such as RT-1, RT-2, and ManipLLM have focused on directly learning robot actions. Such approaches require collecting a significant amount of robot interaction data, which is extremely costly in the real world. Thus, we propose A3VLM, an object-centric, actionable, articulation-aware vision language model. A3VLM focuses on the articulation structure and action affordances of objects. Its representation is robot-agnostic and can be translated into robot actions using simple action primitives. Extensive experiments in both simulation benchmarks and real-world settings demonstrate the effectiveness and stability of A3VLM.

LLMVLMManipulationArticulation
BibTeX
@inproceedings{
huang2024avlm,
title={A3{VLM}: Actionable Articulation-Aware Vision Language Model},
author={Siyuan Huang and Haonan Chang and Yuhan Liu and Yimeng Zhu and Hao Dong and Abdeslam Boularias and Peng Gao and Hongsheng Li},
booktitle={8th Annual Conference on Robot Learning},
year={2024},
url={https://openreview.net/forum?id=lyhS75loxe}
}
A3VLM: Actionable Articulation-Aware Vision Language Model · CoRL 2024