Self-Supervised Learning for Object Pose Estimation Through Active Real Sample Capture
Abstract
6D Object pose estimation is a fundamental component in robotics enabling efficient interaction with the environment. In industrial bin-picking tasks, this problem becomes especially challenging due to difficult object poses, complex occlusions, and inter-object ambiguities. In this work, we propose a novel self-supervised method that automatically collects, labels, and fine-tunes on real images using an eye-in-hand camera setup. We leverage the mobile camera to first obtain reliable ground-truth estimates through multi-view pose estimation, allowing us to subsequently reposition the camera to capture and label real ‘hard case’ samples from the estimated scene. This process enables closure of the sim-to-real gap through large quantities of targeted real training data, generated by comparing differences in model performance between real and synthetically reconstructed scenes and informing the mobile camera on specific poses or areas for data capture. We surpass state-of-the-art performance on a challenging bin-picking benchmark: five out of seven objects surpass a 95% correct detection rate, compared to only one out of seven for previous methods.
BibTeX
@inproceedings{ral2026_selfsupervisedle,
title = {Self-Supervised Learning for Object Pose Estimation Through Active Real Sample Capture},
author = {Alan Li and Angela P. Schoellig},
booktitle = {RA-L 2026},
year = {2026}
}