ICRA 2026poster0 citations

LSV-Loc: LiDAR to Street View Image Crossmodal Localization

Sangmin Lee, Donghyun Choi, Jee-Hwan Ryu

Abstract

Accurate global localization remains a fundamental challenge in autonomous vehicle navigation, especially in previously unexplored areas lacking prior map information. Traditional methods typically rely on high-definition (HD) maps generated through prior traversals or utilize auxiliary sensors such as a global positioning system~(GPS). However, the above approaches are often limited by high costs, scalability issues, and decreased reliability in environments where GPS is unavailable. Moreover, prior methods require that both query and reference data originate from the same sensor modality, restricting to generalize across different sensor types. To address limitations, we propose a novel cross-modal localization framework that enables Light Detection and Ranging~(LiDAR)-equipped vehicles to estimate their global pose by leveraging publicly available Street View images. The proposed method leverages a shared embedding space, learned via a weight-sharing Vision Transformer~(ViT) encoder, to align heterogeneous sensor modalities, specifically LiDAR intensity images and geo-tagged Street View. Shared embedding space enables cross-modal matching for global localization via place recognition, eliminating the need for prior map construction or sensor calibration. Further, to compensate for heading discrepancies between the two modalities, the framework introduces an equirectangular perspective-n-point (PnP) solver by patch-level feature correspondences. Our proposed method enables 3-degree-of-freedom~(DoF) global localization, from a single LiDAR scan and a publicly available Street View. Experiments demonstrate that the proposed method achieves high recall and accurate heading estimation, offering a scalable solution for global localization without r

LocalizationAutonomous Vehicle NavigationRange Sensing