2026
IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human–Robot Interaction
ICRA 2026poster
Vision-Language-Action (VLA) models leverage pretrained vision-language models (VLMs) to couple perception with robotic control, offering a promising path toward general purpose embodied intelligence. However, current SOTA VLAs are primarily pretrained on multimodal tasks with limited relevance to e…