ICRA 20250 citations

A Unified End-to-End Network for Category-Level and Instance-Level Object Pose Estimation from RGB Images

Jiale Ren, Hong Liu, Jinfu Liu, Peifeng Jiang

Abstract

Accurately estimating the 6-DoF pose of objects is a fundamental challenge in computer vision and robotics. While category-level pose estimation based on RGBD data has achieved good performance in recent years, estimating poses solely from RGB images remains a significant challenge. Existing RGB-based category-level methods primarily focus on recovering object point clouds from RGB images, and pose prediction is not performed end-to-end by a network. This paper presents a Category-level and Instance-level Pose Estimation Network (CIPE), which models pose estimation as a set prediction problem and enables direct pose regression from RGB images. To further enhance the network's ability to learn object poses, first, a novel learnable rotation representation that redefines rotation learning within Euclidean space is introduced to facilitate rotation regression. Additionally, we propose a prior-query fusion strategy that utilizes a pre-trained point cloud feature extraction network to integrate categorical object features with bounding boxes, thereby improving the incorporation of category information. Experimental results demonstrate that CIPE significantly outperforms existing RGB-based methods on both category-level and instance-level datasets. The code is available at https://github.com/jialeren/CIPE.

BibTeX
@inproceedings{icra2025_aunifiedendtoend,
  title = {A Unified End-to-End Network for Category-Level and Instance-Level Object Pose Estimation from RGB Images},
  author = {Jiale Ren and Hong Liu and Jinfu Liu and Peifeng Jiang},
  booktitle = {ICRA 2025},
  year = {2025}
}
A Unified End-to-End Network for Category-Level and Instance-Level Object Pose Estimation from RGB Images · ICRA 2025