ICRA 2026poster0 citations

SLoFT: End-To-End Semantic Localization with Floorplan and Transformer

Chaerin Min, Hongsheng Yu, Fengtao Fan, Srinath Sridhar, Qiuxuan Wu, Chao Guo

Abstract

Visual localization is critical for AR navigation, AI-driven audio guidance, and mobile robot localization. How- ever, traditional SLAM methods that rely on pre-built 3D maps suffer from high costs, privacy concerns, and sensitivity to environmental changes. Recent floorplan-based localization methods attempt to addresses these challenges by using 2D floorplans, eliminating the need for 3D map construction. Still, existing approaches are often impractical for real-world applications, as they are limited to specific layouts and fail to generalize beyond their training domains. We propose a novel approach that learns to semantically match visual cues from a camera image to a floorplan image with texts and symbols, inspired by human ability to directly localize oneself using a complex floorplan image. To achieve this, we train a single, unified model on a diverse dataset of 1.2M images and 740K floorplans that we curated, which includes a new collection of semantically-rich, real-world floorplans. This allows our model to generalize effectively to previously unseen areas and demonstrates potential towards zero-shot capabilities. Without making assumptions about camera poses or floorplan structures, our end-to-end model significantly outperforms existing methods and exhibits strong robustness to floorplan rotations, lighting changes, and different camera intrinsics, while effectively leveraging semantic cues like text.

LocalizationVisual LearningDeep Learning for Visual Perception