ICRA 2026poster0 citations

From Dream to Action: Hierarchical Policy Learning with 3D World Imagination for Robotic Manipulation

Wenshuo Wang, Ruiteng Zhao, Tat Joo Teo, Marcelo H Ang Jr, Haiyue Zhu

Abstract

Recent advancements in robotics have focused on developing foundation models capable of generating both actions and future states. Typically, these policies leverage world models to depict human-like imagination. However, most methods remain confined to the 2D domain, where they forecast only the final outcome state rather than the evolving interaction process, thereby offering limited guidance for step-by-step control. To address these limitations, we propose a hierarchical framework that couples 3D imagination, 3D perception, and action generation. A triplane-based world model captures future scene dynamics in a computationally efficient manner, providing predictive cues for decision-making. Based on these representations, the action expert, implemented with a flow-based policy network, converts the outputs of 3D imagination and perception into executable commands. We further introduce an adaptive Classifier-Free Guidance strategy to balance action quality with condition adherence. On Adroit, Meta-World, and real-world tasks, our method achieves a 92% voxel IoU in future state prediction and up to 8% higher success rates than state-of-the-art baselines. The performance gains highlight the effectiveness and generalizability of our method in complex robotic manipulation.

Imitation LearningLearning from DemonstrationDeep Learning in Grasping and Manipulation