DiffuDepGrasp: Diffusion-Based Depth Noise Modeling Empowers Sim-To-Real Robotic Grasping
Yingting Zhou, Wenbo Cui, Weiheng Liu, Guixing Chen, Haoran Li, Dongbin Zhao
Abstract
Accurate spatial-geometric perception remains fundamental to robotic grasping, yet physical artifacts in real depth maps like voids and noise establish a significant sim-to-real gap that critically impedes policy transfer. Training-time strategies like procedural noise injection or learned mappings suffer from data inefficiency due to unrealistic noise simulation, which is often ineffective for grasping tasks that require fine manipulation or dependency on paired datasets heavily. Furthermore, leveraging foundation models to reduce the sim-to-real gap via intermediate representations fails to fully mitigate the domain shift and adds computational overhead during deployment. This work confronts dual challenges of data inefficiency and deployment complexity. We propose DiffuDepGrasp, a deploy-efficient sim-to-real framework enabling zero-shot transfer through simulation-exclusive policy training. Its core innovation, the Diffusion Depth Generator, synthesizes geometrically pristine simulation depth with learned sensor-realistic noise via two synergistic modules. The first Diffusion Depth Module leverages temporal geometric priors to enable sample-efficient training of a conditional diffusion model that captures complex sensor noise distributions, while the second Noise Grafting Module preserves metric accuracy during perceptual artifact injection. Policies trained via our framework require only raw depth inputs during deployment, thus eliminating computational overhead. Extensive sim-to-real validation demonstrates 95.7% average success (SOTA) on 12-object grasping with zero-shot transfer and strong generalization to unseen objects, proving data efficiency and practical value.