2025
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
NeurIPS 2025poster
Recent studies on Vision-Language-Action (VLA) models have shifted from the end-to-end action-generation paradigm toward a pipeline involving task planning followed by action generation, demonstrating improved performance on various complex, long-horizon manipulation tasks. However, existing approac…