GVGNet: Gaze-Directed Visual Grounding for Learning Under-Specified Object Referring Intention
Referring Expression Comprehension (REC) and Referring Expression Segmentation (RES) enable robots to infer human's object referring intention through natural languages. In this letter, Gaze-directed Visual Grounding Network (GVGNet) is proposed to disambiguate human's under-specified object referri