← Search

Wang Zhao

21 accepted papers

2025

DI-PCG: Diffusion-based Efficient Inverse Procedural Content Generation for High-quality 3D Asset Creation

CVPR 2025poster

Procedural Content Generation (PCG) is powerful in creating high-quality 3D contents, yet controlling it to produce desired shapes is difficult and often requires extensive parameter tuning. Inverse Procedural Content Generation aims to automatically find the best parameters under the input conditio…

Cited by 4SourcePDFScholar
2025

DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth Estimation

ICCV 2025poster

Diffusion-based video depth estimation methods have achieved remarkable success with strong generalization ability. However, predicting depth for long videos remains challenging. Existing methods typically split videos into overlapping sliding windows, leading to accumulated scale discrepancies acro…

Cited by 0SourcePDFScholar
2025

StdGEN: Semantic-Decomposed 3D Character Generation from Single Images

CVPR 2025poster

We present StdGEN, an innovative pipeline for generating semantically decomposed high-quality 3D characters from single images, enabling broad applications in virtual reality, gaming, and filmmaking, etc. Unlike previous methods which struggle with limited decomposability, unsatisfactory quality, an…

2025

Tracking Everything in Robotic-Assisted Surgery

ICRA 2025

Accurate tracking of tissues and instruments in videos is crucial for Robotic-Assisted Minimally Invasive Surgery (RAMIS), as it enables the robot to comprehend the surgical scene with precise locations and interactions of tissues and tools. Traditional keypoint-based sparse tracking is limited by f

Cited by 5SourcecodeScholar
2024

AlphaTablets: A Generic Plane Representation for 3D Planar Reconstruction from Monocular Videos

NeurIPS 2024poster

We introduce AlphaTablets, a novel and generic representation of 3D planes that features continuous 3D surface and precise boundary delineation. By representing 3D planes as rectangles with alpha channels, AlphaTablets combine the advantages of current 2D and 3D plane representations, enabling accur…

Cited by 0SourcePDFScholar
2024

Generalizable Thermal-based Depth Estimation via Pre-trained Visual Foundation Model

ICRA 2024poster

Depth estimation is a crucial task in computer vision, applicable to various domains such as 3D reconstruction, robotics, and autonomous driving. In particular, thermal-based depth estimation has unique advantages, including night-time vision. However, the existing depth estimation method remains ch…

Cited by 0SourceScholar
2024

MMPI: a Flexible Radiance Field Representation by Multiple Multi-plane Images Blending

ICRA 2024poster

This paper presents a flexible representation of neural radiance fields based on multi-plane images (MPI), for high-quality view synthesis of complex scenes. MPI with Normalized Device Coordinate (NDC) parameterization is widely used in NeRF learning for its simple definition, easy calculation, and…

Cited by 4SourceScholar
2024

MonoPlane: Exploiting Monocular Geometric Cues for Generalizable 3D Plane Reconstruction

IROS 2024poster

This paper presents a generalizable 3D plane detection and reconstruction framework named MonoPlane. Unlike previous robust estimator-based works (which require multiple images or RGB-D input) and learning-based works (which suffer from domain shift), MonoPlane combines the best of two worlds and es…

Cited by 1SourcecodeScholar
2024

O^2-Recon: Completing 3D Reconstruction of Occluded Objects in the Scene with a Pre-trained 2D Diffusion Model

AAAI 2024technical

Occlusion is a common issue in 3D reconstruction from RGB-D videos, often blocking the complete reconstruction of objects and presenting an ongoing problem. In this paper, we propose a novel framework, empowered by a 2D diffusion-based in-painting model, to reconstruct complete surfaces for the hidd…

2023

DarkFeat: Noise-Robust Feature Detector and Descriptor for Extremely Low-Light RAW Images

AAAI 2023technical

Low-light visual perception, such as SLAM or SfM at night, has received increasing attention, in which keypoint detection and local feature description play an important role. Both handcraft designs and machine learning methods have been widely studied for local feature detection and description, ho…

2022

A Double Branch Next-Best-View Network and Novel Robot System for Active Object Reconstruction

ICRA 2022poster

Next best view (NBV) is a technology that finds the best view sequence for sensor to perform scanning based on partial information, which is the core part for robot active reconstruction. Traditional works are mostly based on the evaluation of candidate views through time-consuming volu-metric trans…

Cited by 14SourceScholar
2022

Deep Reinforcement Learning for Robot Collision Avoidance With Self-State-Attention and Sensor Fusion

RA-L 2022

3D LiDAR sensors can provide 3D point clouds of the environment, and are widely used in automobile navigation; while 2D LiDAR sensors can only provide point cloud in a 2D sweeping plane, and then are only used for navigating robots of small height, e.g., floor mopping robots. In this letter, we prop

Cited by 56SourceScholar
2022

ParticleSfM: Exploiting Dense Point Trajectories for Localizing Moving Cameras in the Wild

ECCV 2022poster

"Estimating the pose of a moving camera from monocular video is a challenging problem, especially due to the presence of moving objects in dynamic environments, where the performance of existing camera pose estimation methods are susceptible to pixels that are not geometrically consistent. To tackle…

2021

A Confidence-Based Iterative Solver of Depths and Surface Normals for Deep Multi-View Stereo

ICCV 2021poster

In this paper, we introduce a deep multi-view stereo (MVS) system that jointly predicts depths, surface normals and per-view confidence maps. The key to our approach is a novel solver that iteratively solves for per-view depth map and normal map by optimizing an energy potential based upon the local…

Cited by 17PDFcodeScholar
2021

NerfingMVS: Guided Optimization of Neural Radiance Fields for Indoor Multi-View Stereo

ICCV 2021poster

In this work, we present a new multi-view depth estimation method that utilizes both conventional SfM reconstruction and learning-based priors over the recently proposed neural radiance fields (NeRF). Unlike existing neural network based optimization method that relies on estimated correspondences,…

Cited by 302PDFcodeScholar
2020

Configuration Space Decomposition for Learning-based Collision Checking in High-DOF Robots

IROS 2020poster

Motion planning for robots of high degrees-of-freedom (DOFs) is an important problem in robotics with sampling-based methods in configuration space \mathcal{C}\mathcal{C} as one popular solution. Recently, machine learning methods have been introduced into sampling-based motion planning methods, whi…

Cited by 8SourceScholar
2020

Towards Better Generalization: Joint Depth-Pose Learning Without PoseNet

CVPR 2020poster

In this work, we tackle the essential problem of scale inconsistency for self supervised joint depth-pose learning. Most existing methods assume that a consistent scale of depth and pose can be learned across all input samples, which makes the learning problem harder, resulting in degraded performan…

Cited by 215PDFcodeScholar
2019

GSPN: Generative Shape Proposal Network for 3D Instance Segmentation in Point Cloud

CVPR 2019poster

We introduce a novel 3D object proposal approach named Generative Shape Proposal Network (GSPN) for instance segmentation in point cloud data. Instead of treating object proposal as a direct bounding box regression problem, we take an analysis-by-synthesis strategy and generate proposals by reconstr…

Cited by 381PDFcodeScholar