2026
GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization
RSS 2026poster
Vision-Language-Action (VLA) models aim for general robot learning by aligning action as a modality within powerful Vision-Language Models (VLM). Existing VLAs rely on end-to-end supervision to implicitly enable the action decoding process to learn task-relevant features. However, without explicit g…