SLoFT: End-To-End Semantic Localization with Floorplan and Transformer
Chaerin Min, Hongsheng Yu, Fengtao Fan, Srinath Sridhar, Qiuxuan Wu, Chao Guo
Abstract
Visual localization is critical for AR navigation, AI-driven audio guidance, and mobile robot localization. How- ever, traditional SLAM methods that rely on pre-built 3D maps suffer from high costs, privacy concerns, and sensitivity to environmental changes. Recent floorplan-based localization methods attempt to addresses these challenges by using 2D floorplans, eliminating the need for 3D map construction. Still, existing approaches are often impractical for real-world applications, as they are limited to specific layouts and fail to generalize beyond their training domains. We propose a novel approach that learns to semantically match visual cues from a camera image to a floorplan image with texts and symbols, inspired by human ability to directly localize oneself using a complex floorplan image. To achieve this, we train a single, unified model on a diverse dataset of 1.2M images and 740K floorplans that we curated, which includes a new collection of semantically-rich, real-world floorplans. This allows our model to generalize effectively to previously unseen areas and demonstrates potential towards zero-shot capabilities. Without making assumptions about camera poses or floorplan structures, our end-to-end model significantly outperforms existing methods and exhibits strong robustness to floorplan rotations, lighting changes, and different camera intrinsics, while effectively leveraging semantic cues like text.