Enhancing 6D Pose Estimation with Cross-modal Fusion Network and Density-peak Keypoint Localization
Current dual-fusion models for 6D pose estimation often lead to increased computational complexity and risk of overfitting with the addition of more networks. To address this, we propose a Cross-modal Fusion Network (CFN), which extracts robust dual-modal features while reducing computation energy a…