2026
WMPO: World Model-based Policy Optimization for Vision-Language-Action Models
ICLR 2026poster
Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation, but their reliance on expert demonstrations limits their ability to learn from failures and perform self-corrections. Reinforcement learning (RL) addresses these through self-improving interact…