ICRA 2026poster0 citations

CMG3D: Compensation towards Modality Gap for Open-Vocabulary Indoor 3D Object Detection

Sheng Zhang, Lian Huai, Yuyu Liu, Xingqun Jiang

Abstract

Open-vocabulary indoor three-dimensional object detection (OVI3DOD) is used to detect any class of objects in indoor scenes with prompts. Owing to the relatively limited three-dimensional (3D) data, most of the OVI3DOD algorithms perform training with pseudo labels transformed from the openvocabulary 2D detection results. For indoor scenes, point clouds are sparse and incomplete. Moreover, there is a gap between different modalities, especially for distant objects. However, existing OVI3DOD algorithms ignore this problem, which weakens the detection performance. Therefore, we propose the Compensation towards Modal Gap for open-vocabulary indoor 3D object detection (CMG3D) approach. CMG3D consists of three modules: multimodal compensation (MC), object proposal filtering (OPF) and pseudo label refinement and generation (PLRG). For the MC, features from images are converted into pseudo voxel space and then summed with the voxel space of the point cloud, which is used to compensate for the modality gap. For the OPF, we filter the object proposals to avoid confusion between the foreground and background. For the PLRG, the predictions from the two-dimensional (2D) detector are refined by the multimodal large language model (LLM) SIGLIP and then transformed to 3D pseudo labels for the training process. Finally, we evaluate CMG3D on two indoor datasets, SUN RGB-D and ScanNet, and achieve state-of-the-art results.

Deep Learning for Visual PerceptionRGB-D PerceptionComputer Vision for Automation
CMG3D: Compensation towards Modality Gap for Open-Vocabulary Indoor 3D Object Detection · ICRA 2026