2026
AR-VLA: Autoregressive Action Expert for Vision–Language–Action Models
RSS 2026poster
We propose a standalone autoregressive (AR) Action Expert that generates actions as a continuous causal sequence while conditioning on refreshable vision-language prefixes. In contrast to existing Vision-Language-Action (VLA) models and diffusion policies that reset temporal context with each new ob…