← Search

John Won

1 accepted papers

2026

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

ICML 2026poster

Augmenting Vision-Language-Action models (VLAs) with world models is promising for robotic policy learning but faces challenges in jointly predicting states and actions due to the modality gap. To address this, we propose DUal-STream diffusion (DUST), a world-model augmented VLA framework featuring …

Cited by 0SourceScholar