Boosting Vehicle-to-Vehicle Collaborative Perception in Bird's-Eye View by Attentive Feature Fusion and Robust Pose Correction
Hong Zhou, Chenyang Lu, Liang Li, Xiangchao Meng, Feng Shao, Qiuping Jiang
Abstract
Collaborative perception enables Connected Autonomous Vehicles (CAVs) to share sensory data, and therefore presents a promising path towards long-range robust environmental understanding by overcoming individual perception limitations such as occlusions. The core challenge of collaborative perception lies in the precise spatial alignment of shared features and their effective fusion. To address this, we propose a novel framework named Effective and Robust Collaborative Perception (ERCP), which is designed to enhance feature fusion with strong robustness against CAV pose errors. Specifically, the robustness is improved by a two-stage coarse-to-fine Pose Correction Module (PCM) that performs feature spatial alignment, and is further boosted by the proposed perturbative training mechanism. Furthermore, a Multi-scale Cross-attention Fusion Module (MCFM) effectively aggregates aligned Bird's-Eye View (BEV) features from multiple CAVs, and leverages cross-attention among different scales to create a comprehensive representation for downstream perception tasks. Extensive evaluations on the OPV2V and V2V4Real datasets demonstrate that ERCP achieves state-of-the-art performance in the collaborative vehicle detection task in BEV. Our code is available at: <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/ZhouH188/ERCP</uri>.
BibTeX
@inproceedings{ral2026_boostingvehiclet,
title = {Boosting Vehicle-to-Vehicle Collaborative Perception in Bird's-Eye View by Attentive Feature Fusion and Robust Pose Correction},
author = {Hong Zhou and Chenyang Lu and Liang Li and Xiangchao Meng and Feng Shao and Qiuping Jiang},
booktitle = {RA-L 2026},
year = {2026}
}