IROS 20250 citations

Task-Guided and Object-Centric Conditioning for Effective and Adaptive Diffusion Policy

Wenshuo Wang, Ruiteng Zhao, Tat Joo Teo, Marcelo H. Ang, Haiyue Zhu

Abstract

Imitation learning has emerged as an effective paradigm for training visuo-motor policies in robotic manipulation. In real-world scenarios, visuo-motor policies are required to be effective, sample-efficient, and capable of adapting to dynamic environments. A key factor influencing these capabilities is the quality of visual representations. Conventional approaches that learn a vision encoder and policy network from scratch often result in suboptimal representations, as the training process tends to prioritize policy optimization over rich semantic feature extraction. Alternatively, while pre-trained large vision models offer strong general-purpose features, they often fail to capture the fine-grained, task-specific information required for effective manipulation. To capture rich and informative visual features, we propose TOC-DP, a novel framework that integrates SlotAttention to facilitate object-centric representation learning. Task-specific segmentation priors are incorporated as an inductive bias to enhance the task-awareness and object-awareness of the learned visual features. The extracted representations are subsequently refined to encode action-aware information during downstream policy learning. Extensive experiments on the Meta-World benchmark and real-world tasks demonstrate that TOC-DP achieves a 30% improvement in success rate over baseline methods during deployment for a variety of scenarios.

BibTeX
@inproceedings{iros2025_taskguidedandobj,
  title = {Task-Guided and Object-Centric Conditioning for Effective and Adaptive Diffusion Policy},
  author = {Wenshuo Wang and Ruiteng Zhao and Tat Joo Teo and Marcelo H. Ang and Haiyue Zhu},
  booktitle = {IROS 2025},
  year = {2025}
}
Task-Guided and Object-Centric Conditioning for Effective and Adaptive Diffusion Policy · IROS 2025