Bridging the Depth Gap: Adaptive Scene-Instance Alignment for Training-Free Depth Refinement in Robotic Manipulation Scenes
Pihong Hou, Yongfang Zhang, Boxuan Ye, Yi Wu
Abstract
Monocular depth estimation foundation models provide robust depth priors with exceptional generalization capabilities; however, their predictions typically lack a reliable metric scale and contain local inconsistencies in a zero-shot setting, which limit their deployment in unconstrained real-world environments. Meanwhile, sensor-based depth measurements are typically sparse or noisy, especially in challenging scenarios such as transparent objects or cluttered robotic scenes. To address this limitation, we propose a training-free adaptive scene-instance alignment framework that refines pseudo-depth maps using sparse ground-truth samples without any additional training or fine-tuning. Our approach integrates three stages of alignment: (1) a global affine transformation via Theil–Sen regression; (2) multi-scale local corrections guided by Gaussian, color-similarity, and gradient-aware weights; and (3) an optional instance-level residual refinement driven by segmentation guidance. Extensive experiments on robotic grasping datasets (Linemod-Occluded, YCB-Video, HOPE, Wild6D) and transparent-object datasets (ClearPose, TransCG) show consistent improvements across evaluation metrics, including root mean square error (RMSE), mean absolute error (MAE), and thresholded accuracy <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$\delta$</tex-math></inline-formula>. The results demonstrate that our method reliably bridges the gap between monocular priors and sensor ground truth, generating pixel-wise metric depth maps with strong structural fidelity that are well-suited for open-world perception tasks.
BibTeX
@inproceedings{ral2026_bridgingthedepth,
title = {Bridging the Depth Gap: Adaptive Scene-Instance Alignment for Training-Free Depth Refinement in Robotic Manipulation Scenes},
author = {Pihong Hou and Yongfang Zhang and Boxuan Ye and Yi Wu},
booktitle = {RA-L 2026},
year = {2026}
}