RA-L 20248 citations

Poses as Queries: End-to-End Image-to-LiDAR Map Localization With Transformers

Jinyu Miao, Kun Jiang, Yunlong Wang, Tuopu Wen, Zhongyang Xiao, Zheng Fu, Mengmeng Yang, Maolin Liu

Abstract

High-precision vehicle localization with commercial setups is a crucial technique for high-level autonomous driving tasks. As a newly emerged approach, monocular localization in LiDAR map achieves promising balance between cost and accuracy, but estimating pose by finding correspondences between such cross-modal sensor data is challenging, thereby damaging the localization accuracy. In this letter, we address the problem by proposing a novel Transformer-based neural network to register 2D images into 3D LiDAR map in an end-to-end manner. We first implicitly represent poses as high-dimensional feature vectors called <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">pose queries</i> and gradually optimize poses by interacting with the retrieved relevant information from cross-modal features using attention mechanism in a proposed POse Estimator Transformer (POET) module. Moreover, we apply a multiple hypotheses aggregation method that estimates the final poses by performing parallel optimization on multiple randomly initialized <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">pose queries</i> to reduce the network uncertainty. Comprehensive analysis and experimental results on public benchmarks conclude that the proposed image-to-LiDAR map localization network could achieve state-of-the-art performances in challenging cross-modal localization tasks.

BibTeX
@inproceedings{ral2024_posesasqueriesen,
  title = {Poses as Queries: End-to-End Image-to-LiDAR Map Localization With Transformers},
  author = {Jinyu Miao and Kun Jiang and Yunlong Wang and Tuopu Wen and Zhongyang Xiao and Zheng Fu and Mengmeng Yang and Maolin Liu and Jin Huang and Zhihua Zhong and Diange Yang},
  booktitle = {RA-L 2024},
  year = {2024}
}