← Search

Zizhuang Wei

5 accepted papers

2026

SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models

CVPR 2026

Large vision-language models (VLMs) show strong multimodal understanding but still struggle with 3D spatial reasoning, such as distance estimation, size comparison, and cross-view consistency. Existing 3D-aware methods either depend on auxiliary 3D information or enhance RGB-only VLMs with geometry

Cited by 0SourceScholar
2024

RPBG: Towards Robust Neural Point-based Graphics in the Wild

ECCV 2024oral

"Point-based representations have recently gained popularity in novel view synthesis, for their unique advantages, , intuitive geometric representation, simple manipulation, and faster convergence. However, based on our observation, these point-based neural re-rendering methods are only expected to…

2021

AA-RMVSNet: Adaptive Aggregation Recurrent Multi-View Stereo Network

ICCV 2021poster

In this paper, we present a novel recurrent multi-view stereo network based on long short-term memory (LSTM) with adaptive aggregation, namely AA-RMVSNet. We firstly introduce an intra-view aggregation module to adaptively extract image features by using context-aware convolution and multi-scale agg…

Cited by 195PDFcodeScholar
2020

Dense Hybrid Recurrent Multi-view Stereo Net with Dynamic Consistency Checking

ECCV 2020poster

In this paper, we propose an efficient and effective dense hybrid recurrent multi-view stereo net with dynamic consistency checking, namely $D^{2}$HC-RMVSNet, for accurate dense point cloud reconstruction. Our novel hybrid recurrent multi-view stereo net consists of two core modules: 1) a light DREN…

2020

Pyramid Multi-view Stereo Net with Self-adaptive View Aggregation

ECCV 2020poster

In this paper, we propose an effective and efficient pyramid multi-view stereo (MVS) net with self-adaptive view aggregation for accurate and complete dense point cloud reconstruction. Different from using mean square variance to generate cost volume in previous deep-learning based MVS methods, our…