Dismantling the Illusion of Vision-Language-Action Models Competence via Explicit Distributional Shifts
Given that simulation can never exhaustively enumerate reality, generalization is the determining factor for whether Vision-Language-Action (VLA) models can translate benchmark success into real-world functionality. However, current evaluation protocols often incentivize mechanical memorization rath…