Refer and Grasp: Vision-Language Guided Continuous Dexterous Grasping
Yayu Huang, Dongxuan Fan, Wen Qi, Daheng Li, Yifan Yang, Yongkang Luo, Jia Sun, Qian Liu
Abstract
Robotic grasping guided by natural language instructions faces challenges due to ambiguities in object descriptions and the need to interpret complex spatial context. Existing visual grounding methods often rely on datasets that fail to capture these complexities, particularly when object categories are vague or undefined. To address these challenges, we make three key contributions. First, we present an automated dataset generation engine for visual grounding in tabletop grasping, combining procedural scene synthesis with template-based referring expression generation, requiring no manual labeling. Second, we introduce the RefGrasp dataset, featuring diverse indoor environments and linguistically challenging expressions for robotic grasping tasks. Third, we propose a visually grounded dexterous grasping framework with continuous grasp generation, validated through extensive real-world robotic experiments. Our work offers a novel approach for language-guided robotic manipulation, providing both a challenging dataset and an effective grasping framework for real-world applications. Project website: https://refer-and-grasp.github.io.
BibTeX
@inproceedings{iros2025_referandgraspvis,
title = {Refer and Grasp: Vision-Language Guided Continuous Dexterous Grasping},
author = {Yayu Huang and Dongxuan Fan and Wen Qi and Daheng Li and Yifan Yang and Yongkang Luo and Jia Sun and Qian Liu and Peng Wang},
booktitle = {IROS 2025},
year = {2025}
}