Multi-View Stereo with Geometric Encoding for Dense Scene Reconstruction
Guidong Yang, Rui Cao, Junjie Wen, Benyun Zhao, Qingxiang Li, Yijun Huang, Lei Lei, Xi Chen
Abstract
Multi-view stereo (MVS) implicitly encodes photometric and geometric cues into the cost volume for multi-view correspondence matching, transferring insufficient geometric cues essential to depth estimation and reconstruction. This paper proposes GE-MVS, a novel multi-view stereo network with geometric encoding for more accurate and complete depth estimation and point cloud reconstruction. First, the cross-view adaptive cost volume aggregation module is proposed to strengthen multi-view geometric cues encoding during cost volume construction. Then, the depth consistency optimization is performed in the 3D point space during learning by invoking ground-truth depth cues from adjacent views. Finally, the surface normal geometries are explicitly encoded to refine the sampled depth hypotheses to be consistent in the local neighbor regions. Extensive experiments on the standard MVS benchmarks including DTU, Tanks and Temples, and BlendedMVS demonstrate the state-of-the-art depth estimation and point cloud reconstruction performance of GE-MVS. The GE-MVS is further deployed in real-world experiments for UAV-based large-scale reconstruction, where our method outperforms the prevalent industrial reconstruction solutions concerning reconstruction efficiency and efficacy. Our project page is: https://cuhk-usr-group.github.io/GE-MVS/
BibTeX
@inproceedings{icra2025_multiviewstereow,
title = {Multi-View Stereo with Geometric Encoding for Dense Scene Reconstruction},
author = {Guidong Yang and Rui Cao and Junjie Wen and Benyun Zhao and Qingxiang Li and Yijun Huang and Lei Lei and Xi Chen and Alan H. F. Lam and Yun-Hui Liu and Ben M. Chen},
booktitle = {ICRA 2025},
year = {2025}
}