2025
Mechanistic Interpretability for Steering Vision-Language-Action Models
CoRL 2025poster
Vision-Language-Action (VLA) models are a promising path to realizing generalist embodied agents that can quickly adapt to new tasks, modalities, and environments. However, methods for interpreting and steering VLAs fall far short of classical robotics pipelines, which are grounded in explicit model…