← Search

Sili Chen

5 accepted papers

2026

Depth Anything 3: Recovering the Visual Space from Any Views

ICLR 2026oral

We present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of minimal modeling, DA3 yields two key insights: a single plain transformer (e.g., vanilla DINOv2 encoder) is sufficient…

Cited by 0SourcecodeScholar
2025

Towards In-the-wild 3D Plane Reconstruction from a Single Image

CVPR 2025highlight

3D plane reconstruction from a single image is a crucial yet challenging topic in 3D computer vision. Previous state-of-the-art (SOTA) methods have focused on training their system on a single dataset from either indoor or outdoor domain, limiting their generalizability across diverse testing data.…

2025

Video Depth Anything: Consistent Depth Estimation for Super-Long Videos

CVPR 2025highlight

Depth Anything has achieved remarkable success in monocular depth estimation with strong generalization ability. However, it suffers from temporal inconsistency in videos, hindering its practical applications. Various methods have been proposed to alleviate this issue by leveraging video generation…

Cited by 12SourcePDFScholar
2024

MonoPlane: Exploiting Monocular Geometric Cues for Generalizable 3D Plane Reconstruction

IROS 2024poster

This paper presents a generalizable 3D plane detection and reconstruction framework named MonoPlane. Unlike previous robust estimator-based works (which require multiple images or RGB-D input) and learning-based works (which suffer from domain shift), MonoPlane combines the best of two worlds and es…

Cited by 1SourcecodeScholar
2018

Noise-Resistant Deep Learning for Object Classification in Three-Dimensional Point Clouds Using a Point Pair Descriptor

RA-L 2018

Object retrieval and classification in point cloud data are challenged by noise, irregular sampling density, and occlusion. To address this issue, we propose a point pair descriptor that is robust to noise and occlusion and achieves high retrieval accuracy. We further show how the proposed descripto

Cited by 21SourceScholar