FROM KNOWING TO DOING PRECISELY: A GENERAL SELF-CORRECTION AND TERMINATION FRAMEWORK FOR VLA MODELS
While vision-language-action (VLA) models for embodied agents integrate perception, reasoning, and control, they remain constrained by two critical weaknesses: first, during grasping tasks, the action tokens generated by the language model often exhibit subtle spatial deviations from the target obje…