Self-Supervised Monocular Depth Estimation from Videos via Pose-Adaptive Reconstruction
Xin Sun, Boqian Liu, Xinchen Ye, Rui Xu, Haojie Li
Abstract
Self-supervised depth estimation from videos involves predicting the depth map of a target frame and the pose changes between source and target frames. The reconstructed source frame is aligned with the target view using the predicted pose and depth information. Precise pose estimation significantly impacts frame reconstruction, which, in turn, affects depth estimation performance. In our approach, we emphasize the importance of pose estimation, often overlooked by existing methods. We propose a self-supervised loss function that improves pose-adaptive reconstruction. Specifically, we decompose the conventional pose network into three parallel branches, each estimating pure translation, pure rotation, and full 6-DoF pose components independently. Our pose-adaptive reconstruction loss selects optimal pose parameterizations to minimize reconstruction errors, mitigating the impact of inaccurate posture. Our proposed pose estimation framework outperforms state-of-the-art methods on benchmark datasets.
BibTeX
@inproceedings{icassp2025_selfsupervisedmo,
title = {Self-Supervised Monocular Depth Estimation from Videos via Pose-Adaptive Reconstruction},
author = {Xin Sun and Boqian Liu and Xinchen Ye and Rui Xu and Haojie Li},
booktitle = {ICASSP 2025},
year = {2025}
}