2026
BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation
RSS 2026poster
Equipping embodied agents with the ability to reason about tasks, foresee physical outcomes, and generate precise actions is essential for general-purpose manipulation. While recent Vision-Language-Action (VLA) models have leveraged pre-trained foundation models, they typically focus on either lingu…