← Search

Hengkai Guo

8 accepted papers

2026

Depth Anything 3: Recovering the Visual Space from Any Views

ICLR 2026oral

We present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of minimal modeling, DA3 yields two key insights: a single plain transformer (e.g., vanilla DINOv2 encoder) is sufficient…

Cited by 0SourcecodeScholar
2025

Towards In-the-wild 3D Plane Reconstruction from a Single Image

CVPR 2025highlight

3D plane reconstruction from a single image is a crucial yet challenging topic in 3D computer vision. Previous state-of-the-art (SOTA) methods have focused on training their system on a single dataset from either indoor or outdoor domain, limiting their generalizability across diverse testing data.…

2025

Video Depth Anything: Consistent Depth Estimation for Super-Long Videos

CVPR 2025highlight

Depth Anything has achieved remarkable success in monocular depth estimation with strong generalization ability. However, it suffers from temporal inconsistency in videos, hindering its practical applications. Various methods have been proposed to alleviate this issue by leveraging video generation…

Cited by 12SourcePDFScholar
2024

MonoPlane: Exploiting Monocular Geometric Cues for Generalizable 3D Plane Reconstruction

IROS 2024poster

This paper presents a generalizable 3D plane detection and reconstruction framework named MonoPlane. Unlike previous robust estimator-based works (which require multiple images or RGB-D input) and learning-based works (which suffer from domain shift), MonoPlane combines the best of two worlds and es…

Cited by 1SourcecodeScholar
2022

ParticleSfM: Exploiting Dense Point Trajectories for Localizing Moving Cameras in the Wild

ECCV 2022poster

"Estimating the pose of a moving camera from monocular video is a challenging problem, especially due to the presence of moving objects in dynamic environments, where the performance of existing camera pose estimation methods are susceptible to pixels that are not geometrically consistent. To tackle…

2021

A Confidence-Based Iterative Solver of Depths and Surface Normals for Deep Multi-View Stereo

ICCV 2021poster

In this paper, we introduce a deep multi-view stereo (MVS) system that jointly predicts depths, surface normals and per-view confidence maps. The key to our approach is a novel solver that iteratively solves for per-view depth map and normal map by optimizing an energy potential based upon the local…

Cited by 17PDFcodeScholar
2020

GPO: Global Plane Optimization for Fast and Accurate Monocular SLAM Initialization

ICRA 2020poster

Initialization is essential to monocular Simultaneous Localization and Mapping (SLAM) problems. This paper focuses on a novel initialization method for monocular SLAM based on planar features. The algorithm starts by homography estimation in a sliding window. It then proceeds to a global plane optim…

Cited by 6SourceScholar