A Multimodal Selective Fusion Approach for Robotic Grasp Detection
Effective fusion of RGB and depth images for robotic grasp detection in complex environments remains a critical challenge. Most existing approaches rely on coarse-grained fusion strategies, such as channel concatenation or simple weighting, which are insufficient to fully capture the complementary n