Adaptive Visual-Tactile Fusion for Contact-Rich Dexterous Manipulation
Zedong Cai, Chuan Liu, Siyuan Li, Peng Liu
Abstract
Effective dexterous manipulation hinges on the dynamic integration of visual context with fine-grained tactile feedback. This remains a significant challenge, as existing methods often rely on static fusion strategies and struggle to learn from isolated tactile features. To address this, we propose a hardware-decoupled tactile representation that learns cross-finger spatial features by training a sparse Transformer on unified tactile images, enabling cross-device generalization. Furthermore, we introduce a tactile-activity-guided adaptive visual-tactile fusion mechanism that dynamically adjusts the influence of vision and touch, showing the contribution of touch feedback upon physical contact. We evaluate our method on a series of contact-rich manipulation tasks requiring fine force control. Experimental results show that our method has an average success rate of over 90%, demonstrating effectiveness compared to other methods. Further analysis shows that adaptive multimodal fusion is essential to complete the dexterous manipulation tasks.
BibTeX
@inproceedings{ral2026_adaptivevisualta,
title = {Adaptive Visual-Tactile Fusion for Contact-Rich Dexterous Manipulation},
author = {Zedong Cai and Chuan Liu and Siyuan Li and Peng Liu},
booktitle = {RA-L 2026},
year = {2026}
}