2026
PV-Ground: Text-Guided Point-Voxel Interaction for 3D Visual Grounding
CVPR 2026
3D visual grounding (VG) aims to localize target objects in 3D scenes based on free-form textual descriptions. Existing 3D VG models predominantly employ point-based backbones for point cloud feature extraction. Such methods require aggressive downsampling of the input point cloud, which sacrifices