Vi-TacMan: Articulated Object Manipulation Via Vision and Touch
Leiyao Cui, Zihang Zhao, Sirui Xie, Wenhuan Zhang, Zhi Han, Yixin Zhu
Abstract
Autonomous manipulation of articulated objects represents a basic skill for robots deployed in human environments. Current vision-based methods can infer object hidden kinematics, but their estimates are sometimes imprecise in driving reliable actions, especially on previously unseen objects. Tactile methods, on the other hand, excel once contact is made, yet they require a reasonable initial guess about where and how to interact. This observation suggests a natural division of labor: vision provides global, coarse guidance, while touch delivers precise, robust execution. Building on this complementarity, we propose a systematic approach, Vi-TacMan, which uses vision to plan and touch to control. We begin by training a vision module that accurately detects holdable and movable parts. Once identified, these parts are then segmented for further processing. From these detections, the system proposes feasible grasps along with a coarse interaction direction modeled by a von Mises-Fisher (vMF) distribution. To enhance directional reasoning, we explicitly incorporate surface normals on movable regions as a geometric prior. This inductive bias clarifies the expected motion and improves generalization to unseen objects, yielding significant gains over baseline methods (all p-values less than 0.0001). Finally, seeded with the vision-derived grasp and motion direction, a tactile-informed controller establishes and maintains stable interactions, enabling reliable execution of the manipulation. Real-world object experiments on diverse objects further confirm reliable manipulation without explicit kinematic models. These findings establish a paradigm for multi-modal robotic perception that could advance autonomous systems operating in complex, unstructured environments.