RA-L 20260 citations

Grasp, Reason, Act: Tactile-Language Model for Zeroshot Sim2real Grasp Stability Prediction and Re-Grasping

Gang Yan, Jiaxian Guo, Khrapchenkov Petr, Zhongchao Zhou, Yaonan Zhu, Yusuke Iwasawa, Yutaka Matsuo

Abstract

Robotic tactile learning is a critical research area for enabling robots to perform complex manipulation tasks with human-like dexterity and adaptability. However, ensuring grasp stability remains one of the most fundamental yet challenging problems in tactile sensing. Existing approaches predominantly focus on classifying stability outcomes, lacking the ability to reason about failure modes or propose corrections. We address this gap by harnessing the inherent reasoning capabilities of Vision-Language Models (VLMs). To do so, we introduce a framework that integrates tactile modality via Quantized Low-Rank Adaptation (QLoRA), enabling parameter-efficient fine-tuning that empowers VLMs to both diagnose grasp instability and generate actionable, natural language corrections. To tackle tactile data scarcity, we develop an automated data collection and annotation pipeline leveraging state-of-the-art tactile sensor simulation, calibrated via a real-to-sim procedure to ensure robust transfer. Extensive experiments across 15 k simulated and 400 real-world trials demonstrate superior zero-shot sim-to-real generalization. In real-time physical evaluations, our approach achieves a 95% grasp success rate on unseen objects, significantly outperforming baselines.

BibTeX
@inproceedings{ral2026_graspreasonactta,
  title = {Grasp, Reason, Act: Tactile-Language Model for Zeroshot Sim2real Grasp Stability Prediction and Re-Grasping},
  author = {Gang Yan and Jiaxian Guo and Khrapchenkov Petr and Zhongchao Zhou and Yaonan Zhu and Yusuke Iwasawa and Yutaka Matsuo},
  booktitle = {RA-L 2026},
  year = {2026}
}