RF-Pose Estimation based on Contrastive Camera-Radar-Images Pretraining
Yen-Hsiang Tseng, Po-Hsuan Tseng
Abstract
This paper proposes a contrastive learning-based pretraining method for using millimeter-wave radar signals for pose estimation. By integrating the camera and radar data in the model pretraining, we first utilize K-means clustering algorithm to divide the data into distinct clusters from the spatial distribution. We propose a contrastive camera-radar-images pertaining (CCRP) technique by conducting contrastive learning on camera-captured coordinates and radar signal heatmaps of the associated cluster. The trained radar heatmap encoder network is based on a vision transformer capable of extracting highly distinctive features, reduces the training difficulty of the subsequent pose estimation network, and improves performance with fine-tuning. Based on the HIBER dataset, the proposed CCRP achieved leading results a marked performance improvement by 31% compared to other unsupervised pertaining methods.
BibTeX
@inproceedings{icassp2025_rfposeestimation,
title = {RF-Pose Estimation based on Contrastive Camera-Radar-Images Pretraining},
author = {Yen-Hsiang Tseng and Po-Hsuan Tseng},
booktitle = {ICASSP 2025},
year = {2025}
}