Enhancing 6D Pose Estimation with Cross-modal Fusion Network and Density-peak Keypoint Localization
Liming Zhang, Qing Li, Zhenhong Chen, Chuan Yan, Xiaojiang Peng
Abstract
Current dual-fusion models for 6D pose estimation often lead to increased computational complexity and risk of overfitting with the addition of more networks. To address this, we propose a Cross-modal Fusion Network (CFN), which extracts robust dual-modal features while reducing computation energy and overfitting risks. The CFN consists of multiple Cross-modal Fusion Modules (CFM), featuring two key components: 1) the Spiking-based Cross-Attention Block (SCA), which utilizes only mask and addition operations, significantly lowering computational energy compared to traditional self-attention; and 2) the Specificity Preserving Block (SPB), designed to mitigate overfitting from multiple CFM layers. Additionally, we introduce a density-peak keypoint localization (DKL) method that resists noise and sparse data, eliminating the need for iterative processes. Extensive experiments on multiple 6D pose estimation benchmarks demonstrate that our CFN method significantly outperforms existing state-of-the-art approaches.
BibTeX
@inproceedings{icassp2025_enhancing6dposee,
title = {Enhancing 6D Pose Estimation with Cross-modal Fusion Network and Density-peak Keypoint Localization},
author = {Liming Zhang and Qing Li and Zhenhong Chen and Chuan Yan and Xiaojiang Peng},
booktitle = {ICASSP 2025},
year = {2025}
}