ICRA 2026poster0 citations

Imagine2Act: Leveraging Object-Action Motion Consistency from Imagined Goals for Robotic Manipulation

Liang Heng, Jiadong Xu, Yiwen Wang, Xiaoqi Li, Muhe Cai, Yan Shen, Juan Zhu, Guanghui Ren

Abstract

Relational object rearrangement (ROR) tasks require a robot to manipulate objects with precise semantic and geometric reasoning. Existing approaches either rely on pre-collected demonstrations that struggle to capture complex geometric constraints, or generate goal-state observations to capture semantic and geometric knowledge but fail to explicitly couple object transformation with action prediction, leading to errors from generative noise. To address these limitations, we propose Imagine2Act, a 3D imitation-learning framework that incorporates semantic and geometric constraints of objects into policy learning to tackle high-precision manipulation tasks. We first generate imagined goal images conditioned on language instructions and reconstruct corresponding 3D point clouds to provide robust semantic and geometric priors. These imagined goal point clouds serve as additional inputs to the policy model, while an object–action consistency strategy with soft pose supervision explicitly aligns predicted action motion with object transformation. This design enables Imagine2Act to reason about object relational goals and achieve accurate, high-precision manipulation across diverse tasks. Experiments in both simulation and real world demonstrate that Imagine2Act outperforms previous state-of-the-art policies.

Deep Learning in Grasping and ManipulationPerception for Grasping and Manipulation
Imagine2Act: Leveraging Object-Action Motion Consistency from Imagined Goals for Robotic Manipulation · ICRA 2026