AAAI 2026technical0 citations

10 Open Challenges Steering the Future of Vision-Language-Action Models

Soujanya Poria, Navonil Majumder, Chia-Yu Hung, Amir Ali Bagherzadeh, Chuan Li, Kenneth Kwok, Ziwei Wang, Cheston Tan

Abstract

Due to their ability of follow natural language instructions, vision-language-action (VLA) models are increasingly preva- lent in the embodied AI arena, following the widespread suc- cess of their precursors—LLMs and VLMs. In this paper, we discuss 10 principal milestones in the ongoing develop- ment of VLA models—multimodality, reasoning, data, eval- uation, cross-robkot action generalization, efficiency, whole- body coordination, safety, agents, and coordination with hu- mans. Furthermore, we discuss the emerging trends of us- ing spatial understanding, modeling world dynamics, post training, and data synthesis—all aiming to reach these mile- stones. Through these discussions, we hope to bring attention to the research avenues that may accelerate the development of VLA models into wider acceptability.

BibTeX
@inproceedings{aaai2026_10openchallenges,
  title = {10 Open Challenges Steering the Future of Vision-Language-Action Models},
  author = {Soujanya Poria and Navonil Majumder and Chia-Yu Hung and Amir Ali Bagherzadeh and Chuan Li and Kenneth Kwok and Ziwei Wang and Cheston Tan and Jiajun Wu and David Hsu},
  booktitle = {AAAI 2026},
  year = {2026}
}