SPE-MVS: Spatial Position Encoding Enhanced Multi-View Stereo with Monocular Depth Priors
Shaoqian Wang, Jiadai Sun, Bosen Hou, Qiang Wang, Bin Fan, Bo Li, Bin Lu, Yuchao Dai
Abstract
Learning-based Multi-View Stereo (MVS) methods have become the mainstream in the field, relying on the construction of cost Learning-based Multi-View Stereo (MVS) methods have become the mainstream in the field, relying on the construction of cost volumes through multi-view feature similarity computation. However, existing methods depend heavily on photometric consistency across views, leading to poor performance in challenging regions. To overcome this limitation, we propose SPE-MVS, a novel MVS framework enhanced with Spatial Position Encoding (SPE). The SPE represents the 3D positional information of pixels in each image within a unified metric space, constructed using monocular depth priors. We integrate the SPE alongside image data as input and introduce a Photometric-Spatial Hybrid Feature Extractor, along with an SPE-enhanced cost volume construction module. These components incorporate spatial position-based similarity computation, substantially improving robustness in challenging areas. Furthermore, we propose a Monocular Depth-guided Enhancement (MDGE) module that enhances depth probability map using monocular depth priors, thereby further boosting the depth estimation performance. Extensive experiments demonstrate that our method significantly improves reconstruction quality in difficult regions and achieves state-of-the-art (SOTA) performance on multiple benchmarks. The code will be released at https://github.com/bdwsq1996/SPE-MVS.
BibTeX
@inproceedings{cvpr2026_spemvsspatialpos,
title = {SPE-MVS: Spatial Position Encoding Enhanced Multi-View Stereo with Monocular Depth Priors},
author = {Shaoqian Wang and Jiadai Sun and Bosen Hou and Qiang Wang and Bin Fan and Bo Li and Bin Lu and Yuchao Dai},
booktitle = {CVPR 2026},
year = {2026}
}