2026
villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
ICLR 2026poster
Vision-Language-Action (VLA) models have emerged as a popular paradigm for learning robot manipulation policies that can follow language instructions and generalize to novel scenarios. Recent works have begun to explore the incorporation of latent actions, abstract representations of motion between…