RA-L 20260 citations

A Multimodal Selective Fusion Approach for Robotic Grasp Detection

Lin Shi, Shuaikang Zhang, Xinwen Zhou, Yuanwei Ma, Ran Wei

Abstract

Effective fusion of RGB and depth images for robotic grasp detection in complex environments remains a critical challenge. Most existing approaches rely on coarse-grained fusion strategies, such as channel concatenation or simple weighting, which are insufficient to fully capture the complementary nature of multimodal information and lack effective multi-scale feature interaction mechanisms. To address these limitations, this letter introduces the multimodal selective fusion network MSF-Net, designed for high-precision grasp detection of diverse objects. Specifically, our proposed Dynamic Modality Attention (DMA) module adaptively highlights the complementary information between low-level RGB and depth features. Concurrently, the Bidirectional Feature Fusion (BFF) module intelligently integrates multi-scale high-level features. Furthermore, the Context-Guided Attention (CGA) module, integrated into the decoding stage, enhances contextual awareness and channel selectivity. This effectively mitigates localization from background clutter and insufficient local features. Extensive experimental evaluations demonstrate that MSF-Net achieves superior performance, reaching grasp detection accuracies of 99.4% on the Cornell dataset and 96.5% on the Jacquard dataset. In addition, evaluations on the challenging CBRGD dataset under cross-scene settings show consistent performance improvements across multiple scenes, further validating the robustness of the proposed method in complex real-world environments. Real-world robotic experiments yield a grasp success rate of 95.3%, demonstrating the reliability and practical applicability of the proposed approach.

BibTeX
@inproceedings{ral2026_amultimodalselec,
  title = {A Multimodal Selective Fusion Approach for Robotic Grasp Detection},
  author = {Lin Shi and Shuaikang Zhang and Xinwen Zhou and Yuanwei Ma and Ran Wei},
  booktitle = {RA-L 2026},
  year = {2026}
}
A Multimodal Selective Fusion Approach for Robotic Grasp Detection · RA-L 2026