Dismantling the Illusion of Vision-Language-Action Models Competence via Explicit Distributional Shifts
Xueyang Zhou, Yangming Xu, Guiyao Tie, Yongchao Chen, Chaoran Hu, Bo Tao, xingwei zhao, Xiang Xiang
Abstract
Given that simulation can never exhaustively enumerate reality, generalization is the determining factor for whether Vision-Language-Action (VLA) models can translate benchmark success into real-world functionality. However, current evaluation protocols often incentivize mechanical memorization rather than robust policy learning, leading to a paradoxical duality of failure: high-scoring models exhibit *spurious invariance* to semantic changes while simultaneously displaying *extreme brittleness* to trivial environmental perturbations. To address this, we introduce **LIBERO-Gen**, a diagnostic benchmark systematically designed to shift evaluation from intuition-driven heuristics to explicit distributional assumptions. Through a hierarchical protocol spanning **In-distribution**, **Compositional**, and **Domain Generalization**, LIBERO-Gen reveals performance stratifications previously masked by standard metrics. Our analysis identifies Pi0.5 as the top performer (64.0% in Spatial-CG; 21.2% in Task-CG). By identifying perceptual instability and action binding collapse as primary failure modes while validating the efficacy of structured ``Stair” sampling, LIBERO-Gen establishes a rigorous baseline for deployment reliability.
BibTeX
@inproceedings{
zhou2026dismantling,
title={Dismantling the Illusion of Vision-Language-Action Models Competence via Explicit Distributional Shifts},
author={Xueyang Zhou and Yangming Xu and Guiyao Tie and Chaoran Hu and Bo Tao and xingwei zhao and Xiang Xiang and Pan Zhou and Lichao Sun and Yongchao Chen},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=R7vl47raSn}
}