2026
RetoVLA: Reusing Register Tokens for Spatial Reasoning in Vision-Language-Action Models
ICRA 2026poster
Vision-Language-Action (VLA) models have demonstrated robust performance across diverse robotic tasks. However, their high memory and computational demands often limit real-time deployment. While existing model compression techniques reduce the parameter footprint, they often drop in 3D spatial reas…