AAAI 2026technical0 citations

Multi-Modal Assistance for Unsupervised Domain Adaptation on Point Cloud 3D Object Detection

Shenao Zhao, Pengpeng Liang, Zhoufan Yang

Abstract

Unsupervised domain adaptation for LiDAR-based 3D object detection (3D UDA) based on the teacher-student architecture with pseudo labels has achieved notable improvements in recent years. Although it is quite popular to collect point clouds and images simultaneously, little attention has been paid to the usefulness of image data in 3D UDA when training the models. In this paper, we propose an approach named MMAssist that improves the performance of 3D UDA with multi-modal assistance. A method is designed to align 3D features between the source domain and the target domain by using image and text features as bridges. More specifically, we project the ground truth labels or pseudo labels to the images to get a set of 2D bounding boxes. For each 2D box, we extract its image feature from a pre-trained vision backbone. A large vision-language model (LVLM) is adopted to extract the box

BibTeX
@inproceedings{aaai2026_multimodalassist,
  title = {Multi-Modal Assistance for Unsupervised Domain Adaptation on Point Cloud 3D Object Detection},
  author = {Shenao Zhao and Pengpeng Liang and Zhoufan Yang},
  booktitle = {AAAI 2026},
  year = {2026}
}
Multi-Modal Assistance for Unsupervised Domain Adaptation on Point Cloud 3D Object Detection · AAAI 2026