RA-L 20250 citations

CMG3D: Compensation Towards Modality Gap for Open-Vocabulary Indoor 3D Object Detection

Sheng Zhang, Lian Huai, Yuyu Liu, Xingqun Jiang

Abstract

For open-vocabulary indoor three-dimensional (3D) object detection (OVI3DOD), there is a gap between the image and the point cloud for indoor scenes, especially on distant objects. However, existing algorithms ignore this problem, which weakens the detection performance. Therefore, we propose Compensation towards the Modal Gap for open-vocabulary indoor 3D object detection (CMG3D). CMG3D consists of three modules: multimodal compensation (MC), object proposal filtering (OPF) and pseudo label refinement and generation (PLRG). In the MC, features from images are converted into the pseudo voxel space and then summed with the voxel space of the point cloud, which is used to compensate for the modality gap, while the OPF filters the object proposals to avoid confusion between the foreground and background. Finally, in the PLRG, the predictions from the two-dimensional (2D) detector are refined by the multimodal large language model (LLM) SigLIP and then transformed into 3D pseudo labels for the training process. Finally, we evaluate CMG3D on two indoor datasets, SUN RGB-D and ScanNet, and achieve state-of-the-art results.

BibTeX
@inproceedings{ral2025_cmg3dcompensatio,
  title = {CMG3D: Compensation Towards Modality Gap for Open-Vocabulary Indoor 3D Object Detection},
  author = {Sheng Zhang and Lian Huai and Yuyu Liu and Xingqun Jiang},
  booktitle = {RA-L 2025},
  year = {2025}
}
CMG3D: Compensation Towards Modality Gap for Open-Vocabulary Indoor 3D Object Detection · RA-L 2025