2026
SVP: Improving Vision-Language-Action Models with Dual Stochastic Visual Prompting
ICRA 2026poster
Vision-Language-Action (VLA) models, such as OpenVLA, hold the promise of generalist robots, yet their performance is often impaired by distracted attention, which we identify as a manifestation of shortcut learning. We posit that the solution lies not in architectural modifications, but in a new tr…