ICASSP 2023accepted0 citations

C2BN: Cross-Modality and Cross-Scale Balance Network for Multi-Modal 3D Object Detection

Bonan Ding, Jin Xie, Jing Nie

Abstract

Multi-modal 3D object detection that classifies and locates objects in 3D space by combining point-clouds captured by lidars and RGB images captured by cameras, serves as the basis for autonomous driving. Most of the existing methods aggregate features from point-clouds and images by plain element-wise additions or multiplications. Although these methods improve detection accuracy, such simple operations have difficulties in balancing both modalities. Further, the multi-level features from images also suffer from imbalance problems in receptive fields. To address the above problems, we propose two novel networks: cross-modality balance network (CMN) and cross-scale balance network (CSN). CMN utilizes cross-modality attention mechanisms to balance the importance and receptive field of two modalities. CSN employs cross-scale attention mechanisms to reduce the imbalance in multi-level features. Experiments are performed on the challenging benchmark: KITTI. The experimental results show consistent improvements in different 3D object detection frameworks, which verifies the effectiveness and generality of our proposed networks.

BibTeX
@inproceedings{icassp2023_c2bncrossmodalit,
  title = {C2BN: Cross-Modality and Cross-Scale Balance Network for Multi-Modal 3D Object Detection},
  author = {Bonan Ding and Jin Xie and Jing Nie},
  booktitle = {ICASSP 2023},
  year = {2023}
}
C2BN: Cross-Modality and Cross-Scale Balance Network for Multi-Modal 3D Object Detection · ICASSP 2023