ICRA 2026poster0 citations

CDV-SLAM: Compact Deep Visual SLAM with Unified Semantic and Geometric Perception

Ya Fan, Rongling Lang

Abstract

Robust monocular visual Simultaneous Localization and Mapping (SLAM) serves as a cornerstone for various applications. However, its performance frequently suffers degradation in challenging scenarios including fast motion, dynamic objects, and scale ambiguity. This paper proposes CDV-SLAM, a compact deep visual SLAM framework that unifies geometric and semantic perception through a shared visual foundation model. A tight semantic-geometric fusion network is devised to predict optical flow in fast motion. Semantic features are efficiently reused to obtain segmentation and monocular depth for dynamic objects exclusion and scale acquisition. To further address scale drift, we introduce local scale correction in bundle adjustment. Experimental results demonstrate a 42% decrease in average Absolute Trajectory Error (ATE) on the KITTI dataset over the state-of-the-art. Furthermore, our flow-only visual odometry surpasses geometric-only methods on the TartanAir and EuRoC datasets, with a marginal speed reduction of 6%. Our code is publicly available at https://github.com/FrankYard/CDV-SLAM.

Deep Learning MethodsSLAMLocalization
CDV-SLAM: Compact Deep Visual SLAM with Unified Semantic and Geometric Perception · ICRA 2026