RA-L 20260 citations

VLAD-Grasp: Vision-Language Adaptive Depth Grasping for Part-Specific Manipulation

Meng Liu, Yanli Zou, Hao Wang, Shiquan Qiu, Fan Lei

Abstract

In household service robot grasping tasks, precisely identifying and grasping specific object parts are essential for ensuring safety and execution efficiency. Although existing Vision–Language Models (VLMs) support open-vocabulary understanding, they still struggle with fine-grained semantic instructions such as “grasp the handle of a cup” or “hold the blade when passing a knife.” They also often fail to identify the correct object when multiple identical items appear in the scene. Conventional 6-DoF grasping approaches also suffer from unstable grasp depth and high collision risk. To overcome these challenges, we construct a new vision–language dataset, OVPGrasping, tailored for fine-grained, task-oriented grasping in household scenarios. We further propose two key modules: a Bidirectional Cross-Modal Enhancement Module (BCEM) that enables mutual enhancement between image and text features, and a Gated Cross-Modal Fusion (GCMF) module that selectively reinforces decoder queries with object-level semantic embeddings. In addition, we design an adaptive 6-DoF depth-aware grasping strategy that improves grasp stability and reduces collision probability. On the OVPGrasping dataset, our method improves detection accuracy by 1.67% over the perception Baseline. In real-robot experiments, the single-object success rate increases from 70.83% to 78.33%, and the cluttered-scene success rate rises from 64% to 72%, validating the effectiveness of our fine-grained semantic understanding and adaptive grasping strategy.

BibTeX
@inproceedings{ral2026_vladgraspvisionl,
  title = {VLAD-Grasp: Vision-Language Adaptive Depth Grasping for Part-Specific Manipulation},
  author = {Meng Liu and Yanli Zou and Hao Wang and Shiquan Qiu and Fan Lei},
  booktitle = {RA-L 2026},
  year = {2026}
}
VLAD-Grasp: Vision-Language Adaptive Depth Grasping for Part-Specific Manipulation · RA-L 2026