NeurIPS 2025poster0 citations

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

Chongkai Gao, Zixuan Liu, Zhenghao Chi, Junshan Huang, Xin Fei, Yiwen Hou, Yuxuan Zhang, Yudi Lin

Abstract

Recent studies on Vision-Language-Action (VLA) models have shifted from the end-to-end action-generation paradigm toward a pipeline involving task planning followed by action generation, demonstrating improved performance on various complex, long-horizon manipulation tasks. However, existing approaches vary significantly in terms of network architectures, planning paradigms, representations, and training data sources, making it challenging for researchers to identify the precise sources of performance gains and determine which component is more difficult to learn. To systematically investigate the impacts of different planning paradigms and representations isolating from network architectures and training data, in this paper, we introduce \name, a unified VLA architecture suite capable of various task planning paradigms, and design a comprehensive suite of controlled experiments across diverse object categories (rigid and deformable), visual modalities (2D and 3D), environments (simulation and real-world), and end-effectors (grippers and dexterous hands). Our results demonstrate that: 1) visually grounded planning representations are generally better than language planning representations; 2) the Hierarchical-VLA paradigm generally achieves superior performance than other paradigms, albeit at the cost of slower training and inference speeds.

Robot ManipulationTask PlanningFoundation Models
BibTeX
@inproceedings{
gao2025vlaos,
title={{VLA}-{OS}: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models},
author={Chongkai Gao and Zixuan Liu and Zhenghao Chi and Junshan Huang and Xin Fei and Yiwen Hou and Yuxuan Zhang and Yudi Lin and Zhirui Fang and Lin Shao},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},
year={2025},
url={https://openreview.net/forum?id=PQYazNKEYo}
}
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models · NeurIPS 2025