CVPR 20260 citations

Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos

Ziren Gong, Xiaohan Li, Fabio Tosi, Jiawei Han, Stefano Mattoccia, Jianfei Cai, Matteo Poggi

Abstract

We present Ov3R, a novel framework for open-vocabulary semantic 3D reconstruction from RGB video streams, designed to advance Spatial AI. The system features two key components: CLIP3R, a CLIP-informed 3D reconstruction module that predicts dense point maps from overlapping clips alongside object-level semantics; and 2D-3D OVS, a 2D-3D open-vocabulary semantic module that lifts 2D features into 3D by learning fused descriptors integrating spatial, geometric, and semantic cues. Unlike prior methods, Ov3R incorporates CLIP semantics directly into the reconstruction process, enabling globally consistent geometry and fine-grained semantic alignment. Our framework achieves state-of-the-art performance in both dense 3D reconstruction and open-vocabulary 3D segmentation.

BibTeX
@inproceedings{cvpr2026_ov3ropenvocabula,
  title = {Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos},
  author = {Ziren Gong and Xiaohan Li and Fabio Tosi and Jiawei Han and Stefano Mattoccia and Jianfei Cai and Matteo Poggi},
  booktitle = {CVPR 2026},
  year = {2026}
}
Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos · CVPR 2026