ICRA 2026poster0 citations

TactEx: An Explainable Multimodal Robotic Interaction Framework for Human-Like Touch and Hardness Estimation

Felix Verstraete, Lan Wei, Wen Fan, Dandan Zhang

Abstract

Accurate perception of object hardness is essential for safe and dexterous contact-rich robotic manipulation. Here, we present TactEx, an explainable multimodal robotic interaction framework that unifies vision, touch, and language for human-like hardness estimation and interactive guidance. While task-agnostic, we demonstrate and evaluate TactEx's capabilities on fruit-ripeness assessment as a representative use case requiring tactile perception and contextual understanding. Our system fuses GelSight-Mini tactile streams with RGB observations and language prompts: a ResNet50 + LSTM estimates hardness from tactile sequential data, and a cross-modal alignment module integrates visual cues with LLM guidance. The resulting interface is explainable and multimodal, enabling users to identify fruit ripeness levels with statistically significant separability (p < 0.01 for all fruit pairs). For touch placement, we compare YOLO with Grounded-SAM (GSAM) and find GSAM to be more robust for fine-grained segmentation and contact-site selection. A lightweight LLM parses user instructions and produces grounded natural-language explanations linked to the tactile outputs. In end-to-end evaluations, TactEx attains 90% task success on simple user queries and generalises to novel tasks without large-scale tuning. These results highlight the promise of combining pretrained visual and tactile models with language grounding to advance explainable, human-like touch perception and decision-making in robotics.

Force and Tactile SensingSensor FusionVisual Servoing
TactEx: An Explainable Multimodal Robotic Interaction Framework for Human-Like Touch and Hardness Estimation · ICRA 2026