RA-L 20239 citations

TransAPR: Absolute Camera Pose Regression With Spatial and Temporal Attention

Chengyu Qiao, Zhiyu Xiang, YuanGang Fan, Tingming Bai, Xijun Zhao, Jingyun Fu

Abstract

Visual relocalization aims to estimate the absolute camera pose from an image or sequential images. Recent works tackle this problem by exploiting deep neural networks to regress camera poses. However, spatial and temporal clues from sequential images still remain underexplored, resulting in inaccurate poses and large outliers. In this letter, we introduce a novel vision Transformer based absolute pose regression model, TransAPR, to tackle this problem. Upon the traditional CNN backbone, we design Transformer based spatial and temporal fusion modules respectively to realize sufficient feature interaction among the neighboring images in the sequence. A hierarchical feature aggregation (HFA) module is further designed to aggregate multi-scale and multi-level features in the pose regressor. Benefiting from these delicate designs, our model is able to generate reliable image representations for absolute pose regression, resulting in more robust localization under challenging environments. We conduct extensive experiments on various indoor and outdoor datasets and show that our method achieves state-of-the-art performance.

BibTeX
@inproceedings{ral2023_transaprabsolute,
  title = {TransAPR: Absolute Camera Pose Regression With Spatial and Temporal Attention},
  author = {Chengyu Qiao and Zhiyu Xiang and YuanGang Fan and Tingming Bai and Xijun Zhao and Jingyun Fu},
  booktitle = {RA-L 2023},
  year = {2023}
}
TransAPR: Absolute Camera Pose Regression With Spatial and Temporal Attention · RA-L 2023