SurgCUT3R: Surgical Scene-Aware Continuous Understanding of Temporal 3D Representation
Kaiyuan Xu, Fangzhou Hong, Daniel Elson, Baoru Huang
Abstract
The reconstruction of surgical scenes from monocular endoscopic video is crucial for advancing robotic-assisted surgery, but applying state-of-the-art general-purpose reconstruction models is hindered by a severe lack of supervised training data and performance degradation over long sequences. To address these challenges, we propose SurgCUT3R, a systematic framework for adapting unified 4D reconstruction models to the surgical domain. Our approach makes three primary contributions. First, we introduce a data generation pipeline that leverages public stereo surgical datasets to create large-scale, metric-scale pseudo-ground-truth depth maps, effectively bridging the data gap. Second, we employ a hybrid supervision strategy that combines our pseudo-GT with geometric self-correction to enhance robustness against inherent data imperfections. Third, we design a hierarchical inference framework that utilizes two specialized models—one for global stability and one for local accuracy—to significantly reduce accumulated pose drift in long videos. Experiments on the public SCARED and StereoMIS datasets demonstrate that our method achieves a highly competitive balance between accuracy and efficiency. It delivers near state-of-the-art pose estimation, offering a practical and effective solution for robust reconstruction in surgical environments.