2026
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
ICML 2026oral
Recent progress in large-scale robotic datasets and vision-language models (VLMs) has advanced research on vision-language-action (VLA) models. However, existing VLA models still face two fundamental challenges: (\textit{i}) producing precise low-level actions from high-dimensional observations, (\t…