ESFUSION: Enhanced LiDAR-camera Fusion Architecture for HD Mapping at Intersection
Abstract
The construction of high-definition (HD) maps at intersections is crucial for autonomous driving and vehicle-to-infrastructure (V2I) collaboration. However, the semantic complexity of intersections poses significant challenges for HD mapping. Previous research has predominantly relied on traditional algorithms to process LiDAR or camera data, which often struggle with occlusion and inherent sensor limitations. To address these challenges, we propose a novel method, called ESFusion for Effective BEV Feature Selection and Fusion. To the best of our knowledge, this is the first work to leverage multi-modal data from intelligent roadside infrastructure, particularly LiDAR and cameras, for generating HD maps at intersections. To enhance multi-modal feature representation in Bird's Eye View (BEV), we design a Cross-modal Channel Exchange (CCE) module that creates multi-scale spatial features and facilitates LiDAR-camera information exchange across channels. Additionally, we introduce a Dynamic Feature Selection (DFS) module to adaptively select the most valuable information between modalities. Comprehensive evaluations on the DAIR-V2X dataset demonstrate that our method outperforms single-modal approaches and existing state-of-the-art fusion methods for vehicle-side applications. Moreover, experiments on the nuScenes dataset further highlight the high flexibility of our proposed module, showcasing its ability to be seamlessly integrated into existing multi-modal fusion workflows.
BibTeX
@inproceedings{iros2025_esfusionenhanced,
title = {ESFUSION: Enhanced LiDAR-camera Fusion Architecture for HD Mapping at Intersection},
author = {Suhui Yang and Jingjing Cui},
booktitle = {IROS 2025},
year = {2025}
}