2025
Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
CoRL 2025poster
Vision-Language-Action (VLA) models offer a pivotal approach to learning robotic manipulation at scale by repurposing large pre-trained Vision-Language-Models (VLM) to output robotic actions. However, adapting VLMs for robotic domains comes with an unnecessarily high computational cost, which we att…