GeoLanG: Geometry-Aware Language-Guided Grasping with Unified RGB-D Multimodal Learning
Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains highly challenging in cluttered or occluded scenes. Existing methods often rely on multi-stage pipelines that separate obje…