2026
Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
RSS 2026poster
Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-specific action supervision. Since these models typically rely on VLM backbones optimized for Visual Question Answering (VQA…