ICRA 2026poster0 citations

SVP: Improving Vision-Language-Action Models with Dual Stochastic Visual Prompting

Zhide Zhong, Haodong Yan, Tianran Zhang, Lujia Wang, Jin Wu, Jun Ma, Xinhu Zheng, Haoang Li

Abstract

Vision-Language-Action (VLA) models, such as OpenVLA, hold the promise of generalist robots, yet their performance is often impaired by distracted attention, which we identify as a manifestation of shortcut learning. We posit that the solution lies not in architectural modifications, but in a new training paradigm centered on visual prompts that provide explicit visual guidance to the model. We introduce Dual Stochastic Visual Prompting (SVP) as a concrete realization of this paradigm. SVP functions as a training-only ``visual scaffold'', a non-invasive mechanism that requires no architectural modifications. Our work demonstrates that this data-centric training paradigm is a highly effective strategy for mitigating distracted attention, enabling the learning of more robust and capable policies without architectural overhead. SVP yields substantial gains on the challenging LIBERO benchmark and real robot experience. It improves the absolute success rate of the standard OpenVLA by 8.2% on long-horizon tasks and enhances the performance of the highly optimized OpenVLA-OFT. These improvements are validated on a real robot, where our model consistently outperforms baselines across a variety of manipulation tasks.

Imitation LearningRepresentation LearningMachine Learning for Robot Control