PGCSPose: Physics-Constrained Generation and Causal Semantic Fusion for Robust In-Hand Pose Estimation
Peiliang Wu, Yao Li, Mingyue Niu, Wenbai Chen, Guowei Gao
Abstract
Accurate in-hand pose estimation is essential for dexterous robotic manipulation but remains fragile under severe visual occlusion (<inline-formula><tex-math notation="LaTeX">$>$</tex-math></inline-formula>50%) and intermittent tactile contact. Existing visuo-tactile fusion methods treat vision and touch symmetrically, overlooking the asymmetric dependency: visual observations (object geometry and grasp) influence tactile feedback through contact mechanics. We present <bold>PGCSPose</bold>, a framework that leverages this vision<inline-formula><tex-math notation="LaTeX">$\rightarrow$</tex-math></inline-formula>tactile dependency via two components: <italic>TacDiffusion</italic>, a physics-constrained diffusion model that synthesizes plausible tactile signals under contact loss, and <italic>CauCon</italic>, a reasoning module inspired by causal principles that integrates semantic priors from language models for robust inference. Extensive evaluation on ObjectInHand and VinT-6D demonstrates that PGCSPose reduces <bold>pose error by <inline-formula><tex-math notation="LaTeX">$>$</tex-math></inline-formula>50%</bold> under 70% occlusion and 60% tactile dropout, achieving position errors of 0.44–0.47 cm and angular errors of 0.075–0.081 rad across benchmarks, while maintaining real-time performance (16.1 FPS). These results demonstrate that physics-constrained generation and causally-inspired semantic reasoning enable reliable manipulation under degraded sensing conditions.
BibTeX
@inproceedings{ral2026_pgcsposephysicsc,
title = {PGCSPose: Physics-Constrained Generation and Causal Semantic Fusion for Robust In-Hand Pose Estimation},
author = {Peiliang Wu and Yao Li and Mingyue Niu and Wenbai Chen and Guowei Gao},
booktitle = {RA-L 2026},
year = {2026}
}