A 3D Vision-Based Framework for Teleoperation and Dynamic Catching With a High-Speed Multi-Fingered Hand
Xiaohang Shi, Qitong Guo, Ruoyu Jia, Chunxin Yang, Kenichi Murakami, Yuji Yamakawa
Abstract
Driven by significant advancements in structure, sensors, and control algorithms, multi-fingered hand systems have received increasing attention from both academia and industry. While existing research primarily focuses on the interaction static rigid objects manipulation, real-world scenarios often involve dynamic tasks and non-rigid materials. To address these challenges, we developed a multi-fingered manipulation system designed to catch dynamic, non-rigid objects with small-deformation. We decomposed this objective into three components: teleoperation, reinforcement learning-based control, and deformation feedback. This paper presents a 3D vision-based framework that integrates these elements. Experimental results demonstrate that the framework can catch a paper sheet with 5.40 mm deformation and a paper balloon with 12.19 mm deformation. Regarding system latency, the latencies for teleoperation, RL-based control, and deformation feedback are 1.29 ms, 0.65 ms, and 0.47 ms, respectively.
BibTeX
@inproceedings{ral2026_a3dvisionbasedfr,
title = {A 3D Vision-Based Framework for Teleoperation and Dynamic Catching With a High-Speed Multi-Fingered Hand},
author = {Xiaohang Shi and Qitong Guo and Ruoyu Jia and Chunxin Yang and Kenichi Murakami and Yuji Yamakawa},
booktitle = {RA-L 2026},
year = {2026}
}