MASTD3R-SLAM: Monocular Adaptive Semantic Tracking and Dynamic Reconstruction SLAM
Fengwei Yang, Qingran Lin, Chaolun Zhu
Abstract
The challenge of dynamic scenes has long been one of the core issues in the application and generalization of SLAM systems. Traditional visual SLAM systems often rely on depth sensors and prior camera parameters, making it difficult to correct dynamic challenges from arbitrary input images while simultaneously constructing dense maps. Recently, neural network-based methods for two-view point cloud prediction have gained attention, and SLAM systems such as DUST3R and MAST3R have emerged based on this approach. However, these systems face challenges when applied to dynamic scenes and cannot directly use traditional methods for correction, such as semantic masking or optical flow segmentation. To address this issue, we propose MASTD3R-SLAM, a SLAM method specifically designed for dynamic scenes that supports arbitrary video inputs. The method combines fused mask-based processing with coarse-to-fine pointmap alignment and optimization to achieve point cloud–to–pose re-mapping correction, and further performs Gaussian rendering to remove rendering artifacts and suppress dynamic mapping interference. Compared to the original baseline, our approach improves tracking ATE accuracy by more than 90% and successfully restores the correct 3D map.