2024
ViewInfer3D: 3D Visual Grounding Based on Embodied Viewpoint Inference
RA-L 2024
3D Visual Grounding (3D VG) is a fundamental task in embodied intelligence, which entails robots interpreting natural language descriptions to locate objects within 3D environments. The complexity of this task emerges as robots perceive the spatial relationships of objects differently depending on t