← Search

Dave Zhenyu Chen

9 accepted papers

2026

AnchorSplat: Feed-Forward 3D Gaussian Splatting With 3D Geometric Priors

CVPR 2026

Recent feed-forward Gaussian reconstruction models adopt a pixel-aligned formulation that maps each 2D pixel to a 3D Gaussian, entangling Gaussian representations tightly with the input images. In this paper, we propose AnchorSplat, a novel feed-forward 3DGS framework for scene-level reconstruction

Cited by 0SourceScholar
2026

Reliev3R: Relieving Feed-forward 3D Reconstruction from Multi-View Geometric Annotations

CVPR 2026

With recent advances, Feed-forward Reconstruction Models (FFRMs) have demonstrated great potential in reconstruction quality and adaptiveness to multiple downstream tasks. However, the excessive reliance on multi-view geometric annotations, e.g. 3D point maps and camera poses, makes the fully-superv

Cited by 0SourceScholar
2026

VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection

CVPR 2026

Current multi-view indoor 3D object detectors rely on sensor geometry that is costly to obtain--i.e., precisely calibrated multi-view camera poses--to fuse multi-view information into a global scene representation, limiting deployment in real-world scenes. We target a more practical setting: Sensor-

Cited by 0SourcecodeScholar
2026

WPT: World-to-Policy Transfer via Online World Model Distillation

CVPR 2026

Recent years have witnessed remarkable progress in world models, which primarily aim to capture the spatiotemporal correlations between an agent's actions and the evolving environment. However, existing approaches often suffer from tight runtime coupling or depend on offline reward signals, resultin

Cited by 0SourceScholar
2025

Taming Video Diffusion Prior with Scene-Grounding Guidance for 3D Gaussian Splatting from Sparse Inputs

CVPR 2025highlight

Despite recent successes in novel view synthesis using 3D Gaussian Splatting (3DGS), modeling scenes with sparse inputs remains a challenge. In this work, we address two critical yet overlooked issues in real-world sparse-input modeling: extrapolation and occlusion. To tackle these issues, we propos…

Cited by 0SourcePDFScholar
2024

EchoScene: Indoor Scene Generation via Information Echo over Scene Graph Diffusion

ECCV 2024poster

"We present EchoScene, an interactive and controllable generative model that generates 3D indoor scenes on scene graphs. EchoScene leverages a dual-branch diffusion model that dynamically adapts to scene graphs. Existing methods struggle to handle scene graphs due to varying numbers of nodes, multip…

Cited by 23SourcePDFScholar
2024

SceneTex: High-Quality Texture Synthesis for Indoor Scenes via Diffusion Priors

CVPR 2024highlight

We propose SceneTex a novel method for effectively generating high-quality and style-consistent textures for indoor scenes using depth-to-image diffusion priors. Unlike previous methods that either iteratively warp 2D views onto a mesh surface or distillate diffusion latent features without accurate…

Cited by 30SourcePDFScholar
2023

Text2Tex: Text-driven Texture Synthesis via Diffusion Models

ICCV 2023poster

We present Text2Tex, a novel method for generating high-quality textures for 3D meshes from the given text prompts. Our method incorporates inpainting into a pre-trained depth-aware image diffusion model to progressively synthesize high resolution partial textures from multiple viewpoints. To avoid…

Cited by 179PDFScholar
2020

ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language

ECCV 2020poster

We introduce the new task of 3D object localization in RGB-D scans using natural language descriptions. As input, we assume a point cloud of a scanned 3D scene along with a free-form description of a specified target object. To address this task, we propose ScanRefer, where the core idea is to learn…

Cited by 413SourcePDFScholar