← Search

Kefan Gu

1 accepted papers

2026

IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human–Robot Interaction

ICRA 2026poster

Vision-Language-Action (VLA) models leverage pretrained vision-language models (VLMs) to couple perception with robotic control, offering a promising path toward general purpose embodied intelligence. However, current SOTA VLAs are primarily pretrained on multimodal tasks with limited relevance to e…