← Search

Zhaopeng Cui

64 accepted papers

2026

Archon: A Unified Multimodal Model for Holistic Digital Human Generation

CVPR 2026

Digital humans are fundamental to immersive interaction, yet creating a unified model for holistic modalities, including text, audio, motion, and visual content, remains an open challenge. In this paper, we present Archon, a fully pretrained, human-centric unified multimodal model for holistic avata

Cited by 0SourceScholar
2026

DiffWind: Physics-Informed Differentiable Modeling of Wind-Driven Object Dynamics

ICLR 2026poster

Modeling wind-driven object dynamics from video observations is highly challenging due to the invisibility and spatio–temporal variability of wind, as well as the complex deformations of objects. We present DiffWind, a physics-informed differentiable framework that unifies wind–object interaction mo…

Cited by 0SourcecodeScholar
2026

GaussianZoom: Progressive Zoom-in Generative 3D Gaussian Splatting with Geometric and Semantic Guidance

CVPR 2026

We introduce GaussianZoom, a generative zoom-in 3D reconstruction system with an iterative progressive framework that combines geometry-consistent scene modeling and multi-scale semantic reasoning to enable high-fidelity extreme zoom-in rendering from low-resolution inputs. To achieve this, we devel

Cited by 0SourceScholar
2026

PhysSkin: Real-Time and Generalizable Physics-Based Animation via Self-Supervised Neural Skinning

CVPR 2026

Achieving real-time physics-based animation that generalizes across diverse 3D shapes and discretizations remains a fundamental challenge. We introduce PhysSkin, a physics-informed framework that addresses this challenge. In the spirit of Linear Blend Skinning, we learn continuous skinning fields as

Cited by 0SourcecodeScholar
2025

AccidentalGS: 3D Gaussian Splatting from Accidental Camera Motion

ICCV 2025poster

Neural 3D modeling and novel view synthesis with Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS) typically requires the multi-view images with wide baselines and accurate camera poses as input. However, scenarios with accidental camera motions are rarely studied. In this paper, we prop…

Cited by 0SourcePDFScholar
2025

AtlasGS: Atlanta-world Guided Surface Reconstruction with Implicit Structured Gaussians

NeurIPS 2025poster

3D reconstruction of indoor and urban environments is a prominent research topic with various downstream applications. However, existing geometric priors for addressing low-texture regions in indoor and urban settings often lack global consistency. Moreover, Gaussian Splatting and implicit SDF fiel…

Cited by 0SourcecodeScholar
2025

BlinkTrack: Feature Tracking over 80 FPS via Events and Images

ICCV 2025poster

Event cameras, known for their high temporal resolution and ability to capture asynchronous changes, have gained significant attention for their potential in feature tracking, especially in challenging conditions. However, event cameras lack the fine-grained texture information that conventional cam…

2025

Free360: Layered Gaussian Splatting for Unbounded 360-Degree View Synthesis from Extremely Sparse and Unposed Views

CVPR 2025poster

Neural rendering has demonstrated remarkable success in high-quality 3D neural reconstruction and novel view synthesis with dense input views and accurate poses. However, applying it to sparse, unposed views in unbounded 360* scenes remains a challenging problem. In this paper, we propose a novel ne…

2025

GURecon: Learning Detailed 3D Geometric Uncertainties for Neural Surface Reconstruction

AAAI 2025technical

Neural surface representation has demonstrated remarkable success in the areas of novel view synthesis and 3D reconstruction. However, assessing the geometric quality of 3D reconstructions in the absence of ground truth mesh remains a significant challenge, due to its rendering-based optimization pr…

Cited by 0SourcePDFScholar
2025

GaussianUpdate: Continual 3D Gaussian Splatting Update for Changing Environments

ICCV 2025poster

Novel view synthesis with neural models has advanced rapidly in recent years, yet adapting these models to scene changes remains an open problem. Existing methods are either labor-intensive, requiring extensive model retraining, or fail to capture detailed types of changes over time. In this paper,…

Cited by 0SourcePDFScholar
2025

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular VideosC

CVPR 2025poster

We propose a novel 3D-aware diffusion-based method for generating photorealistic talking head videos directly from a single identity image and explicit control signals (e.g., expressions). Our method generates Multiplane Images (MPIs) that ensure geometric consistency, making them ideal for immersiv…

Cited by 0SourcePDFScholar
2025

InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction from Cluttered Scenes

ICCV 2025poster

Humans can naturally identify and mentally complete occluded objects in cluttered environments. However, imparting similar cognitive ability to robotics remains challenging even with advanced reconstruction techniques, which models scenes as undifferentiated wholes and fails to recognize complete ob…

Cited by 0SourcePDFScholar
2025

LightCity: An Urban Dataset for Outdoor Inverse Rendering and Reconstruction under Multi-illumination Conditions

ICCV 2025poster

We propose an outdoor scene dataset and propose a series of benchmarks based on it.Inverse rendering in urban scenes is pivotal for applications like autonomous driving and digital twins, yet it faces significant challenges due to complex illumination conditions, including multi-illumination and ind…

Cited by 0SourcePDFScholar
2025

Neuraloc: Visual Localization in Neural Implicit Map With Dual Complementary Features

ICRA 2025

Recently, neural radiance fields (NeRF) have gained significant attention in the field of visual localization. However, existing NeRF-based approaches either lack geometric constraints or require extensive storage for feature matching, limiting their practical applications. To address these challeng

Cited by 6SourcecodeScholar
2024

"BlinkVision: A Benchmark for Optical Flow, Scene Flow and Point Tracking Estimation using RGB Frames and Events"

ECCV 2024poster

"Recent advances in event-based vision suggest that they complement traditional cameras by providing continuous observation without frame rate limitations and high dynamic range which are well-suited for correspondence tasks such as optical flow and point tracking. However, so far there is still a l…

Cited by 4SourcePDFScholar
2024

A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding

NeurIPS 2024poster

In this paper, we propose a novel multi-view stereo (MVS) framework that gets rid of the depth range prior. Unlike recent prior-free MVS methods that work in a pair-wise manner, our method simultaneously considers all the source images. Specifically, we introduce a Multi-view Disparity Attention (MD…

Cited by 0SourcePDFScholar
2024

GeneAvatar: Generic Expression-Aware Volumetric Head Avatar Editing from a Single Image

CVPR 2024poster

Recently we have witnessed the explosive growth of various volumetric representations in modeling animatable head avatars. However due to the diversity of frameworks there is no practical method to support high-level applications like 3D head avatar editing across different representations. In this…

2024

PNeRFLoc: Visual Localization with Point-Based Neural Radiance Fields

AAAI 2024technical

Due to the ability to synthesize high-quality novel views, Neural Radiance Fields (NeRF) has been recently exploited to improve visual localization in a known environment. However, the existing methods mostly utilize NeRF for data augmentation to improve the regression model training, and their perf…

2023

BlinkFlow: A Dataset to Push the Limits of Event-Based Optical Flow Estimation

IROS 2023poster

Event cameras provide high temporal precision, low data rates, and high dynamic range visual perception, which are well-suited for optical flow estimation. While data-driven optical flow estimation has obtained great success in RGB cameras, its generalization performance is seriously hindered in eve…

Cited by 38SourcecodeScholar
2023

CP-SLAM: Collaborative Neural Point-based SLAM System

NeurIPS 2023poster

This paper presents a collaborative implicit neural simultaneous localization and mapping (SLAM) system with RGB-D image sequences, which consists of complete front-end and back-end modules including odometry, loop detection, sub-map fusion, and global refinement. In order to enable all these module…

Cited by 29SourcePDFScholar
2023

DPS-Net: Deep Polarimetric Stereo Depth Estimation

ICCV 2023poster

Stereo depth estimation usually struggles to deal with textureless scenes for both traditional and learning-based methods due to the inherent dependence on image correspondence matching. In this paper, we propose a novel neural network, i.e., DPS-Net, to exploit both the prior geometric knowledge an…

Cited by 28PDFScholar
2023

IntrinsicNeRF: Learning Intrinsic Neural Radiance Fields for Editable Novel View Synthesis

ICCV 2023poster

Existing inverse rendering combined with neural rendering methods can only perform editable novel view synthesis on object-specific scenes, while we present intrinsic neural radiance fields, dubbed IntrinsicNeRF, which introduce intrinsic decomposition into the NeRF-based neural rendering method and…

Cited by 58PDFcodeScholar
2023

Multi-Modal Neural Radiance Field for Monocular Dense SLAM with a Light-Weight ToF Sensor

ICCV 2023poster

Light-weight time-of-flight (ToF) depth sensors are compact and cost-efficient, and thus widely used on mobile devices for tasks such as autofocus and obstacle detection. However, due to the sparse and noisy depth measurements, these sensors have rarely been considered for dense geometry reconstruct…

Cited by 32PDFcodeScholar
2023

Novel-View Synthesis and Pose Estimation for Hand-Object Interaction from Sparse Views

ICCV 2023poster

Hand-object interaction understanding and the barely addressed novel view synthesis are highly desired in the immersive communication, whereas it is challenging due to the high deformation of hand and heavy occlusions between hand and object. In this paper, we propose a neural rendering and pose est…

Cited by 16PDFcodeScholar
2023

PATS: Patch Area Transportation With Subdivision for Local Feature Matching

CVPR 2023poster

Local feature matching aims at establishing sparse correspondences between a pair of images. Recently, detector-free methods present generally better performance but are not satisfactory in image pairs with large scale differences. In this paper, we propose Patch Area Transportation with Subdivision…

Cited by 42SourcePDFScholar
2023

SINE: Semantic-Driven Image-Based NeRF Editing With Prior-Guided Editing Field

CVPR 2023poster

Despite the great success in 2D editing using user-friendly tools, such as Photoshop, semantic strokes, or even text prompts, similar capabilities in 3D areas are still limited, either relying on 3D modeling skills or allowing editing within only a few categories. In this paper, we present a novel s…

2022

CompNVS: Novel View Synthesis with Scene Completion

ECCV 2022poster

"We introduce a scalable framework for novel view synthesis from RGB-D images with largely incomplete scene coverage. While generative neural approaches have demonstrated spectacular results on 2D images, they have not yet achieved similar photorealistic results in combination with scene completion…

Cited by 8SourcePDFScholar
2022

Crossview Mapping with Graph-based Geolocalization on City-Scale Street Maps

ICRA 2022poster

3D environment mapping has been actively stud-ied recently with the development of autonomous driving and augmented reality. Although many image-based methods are proposed due to their convenience and flexibility compared to other complex sensors, few works focus on fixing the inherent scale ambigui…

Cited by 5SourceScholar
2022

DELTAR: Depth Estimation from a Light-Weight ToF Sensor and RGB Image

ECCV 2022poster

"Light-weight time-of-flight (ToF) depth sensors are small, cheap, low-energy and have been massively deployed on mobile devices for the purposes like autofocus, obstacle detection, etc. However, due to their specific measurements (depth distribution in a region instead of the depth value at a certa…

2022

Generative Category-Level Shape and Pose Estimation with Semantic Primitives

CoRL 2022poster

Empowering autonomous agents with 3D understanding for daily objects is a grand challenge in robotics applications. When exploring in an unknown environment, existing methods for object pose estimation are still not satisfactory due to the diversity of object shapes. In this paper, we propose a nove…

Cited by 28SourcecodeScholar
2022

NICE-SLAM: Neural Implicit Scalable Encoding for SLAM

CVPR 2022poster

Neural implicit representations have recently shown encouraging results in various domains, including promising progress in simultaneous localization and mapping (SLAM). Nevertheless, existing methods produce over-smoothed scene reconstructions and have difficulty scaling up to large scenes. These l…

Cited by 766PDFcodeScholar
2022

NeuMesh: Learning Disentangled Neural Mesh-Based Implicit Field for Geometry and Texture Editing

ECCV 2022poster

"Very recently neural implicit rendering techniques have been rapidly evolved and shown great advantages in novel view synthesis and 3D scene reconstruction. However, existing neural rendering methods for editing purposes offer limited functionality, e.g., rigid transformation, or not applicable for…

2022

SceneSqueezer: Learning To Compress Scene for Camera Relocalization

CVPR 2022oral

Standard visual localization methods build a priori 3D model of a scene which is used to establish correspondences against the 2D keypoints in a query image. Storing these pre-built 3D scene models can be prohibitively expensive for large-scale environments, especially on mobile devices with limited…

Cited by 37PDFScholar
2021

Coxgraph: Multi-Robot Collaborative, Globally Consistent, Online Dense Reconstruction System

IROS 2021poster

Real-time dense reconstruction has been extensively studied for its wide applications in computer vision and robotics, meanwhile much effort has been made for the multi-robot system which plays an irreplaceable role in complicated but time-critical scenarios, e.g., search and rescue tasks. In this p…

Cited by 14SourceScholar
2021

DeepPanoContext: Panoramic 3D Scene Understanding With Holistic Scene Context Graph and Relation-Based Optimization

ICCV 2021poster

Panorama images have a much larger field-of-view thus naturally encode enriched scene context information compared to standard perspective images, which however is not well exploited in the previous scene understanding methods. In this paper, we propose a novel method for panoramic 3D scene understa…

Cited by 40PDFcodeScholar
2021

End-to-End Rotation Averaging With Multi-Source Propagation

CVPR 2021poster

This paper presents an end-to-end neural network for multiple rotation averaging in SfM. Due to the manifold constraint of rotations, conventional methods usually take two separate steps involving spanning tree based initialization and iterative nonlinear optimization respectively. These methods can…

Cited by 29PDFcodeScholar
2021

Graph-Based Asynchronous Event Processing for Rapid Object Recognition

ICCV 2021poster

Different from traditional video cameras, event cameras capture asynchronous events stream in which each event encodes pixel location, trigger time, and the polarity of the brightness changes. In this paper, we introduce a novel graph-based framework for event cameras, namely SlideGCN. Unlike some r…

Cited by 110PDFScholar
2021

Holistic 3D Scene Understanding From a Single Image With Implicit Representation

CVPR 2021poster

We present a new pipeline for holistic 3D scene understanding from a single image, which could predict object shape, object pose and scene layout. As it is a highly ill-posed problem, existing methods usually suffer from inaccurate estimation of both shapes and layout especially for the cluttered sc…

Cited by 129PDFcodeScholar
2021

Learning Object-Compositional Neural Radiance Field for Editable Scene Rendering

ICCV 2021poster

Implicit neural rendering techniques have shown promising results for novel view synthesis. However, existing methods usually encode the entire scene as a whole, which is generally not aware of the object identity and limits the ability to the high-level editing tasks such as moving or adding furnit…

Cited by 356PDFScholar
2021

P2-Net: Joint Description and Detection of Local Features for Pixel and Point Matching

ICCV 2021poster

Accurately describing and detecting 2D and 3D keypoints is crucial to establishing correspondences across images and point clouds. Despite a plethora of learning-based 2D or 3D local feature descriptors and detectors having been proposed, the derivation of a shared descriptor and joint keypoint dete…

Cited by 62PDFcodeScholar
2021

Sat2Vid: Street-View Panoramic Video Synthesis From a Single Satellite Image

ICCV 2021poster

We present a novel method for synthesizing both temporally and geometrically consistent street-view panoramic video from a single satellite image and camera trajectory. Existing cross-view synthesis approaches focus on images, while video synthesis in such a case has not yet received enough attentio…

Cited by 12PDFScholar
2021

Towards Efficient Graph Convolutional Networks for Point Cloud Handling

ICCV 2021poster

We aim at improving the computational efficiency of graph convolutional networks (GCNs) for learning on point clouds. The basic graph convolution that is composed of a K-nearest neighbor (KNN) search and a multilayer perceptron (MLP) is examined. By mathematically analyzing the operations there, two…

Cited by 34PDFScholar
2021

Vis2Mesh: Efficient Mesh Reconstruction From Unstructured Point Clouds of Large Scenes With Learned Virtual View Visibility

ICCV 2021poster

We present a novel framework for mesh reconstruction from unstructured point clouds by taking advantage of the learned visibility of the 3D points in the virtual views and traditional graph-cut based mesh generation. Specifically, we first propose a three-step network that explicitly employs depth c…

Cited by 17PDFcodeScholar
2020

DIST: Rendering Deep Implicit Signed Distance Function With Differentiable Sphere Tracing

CVPR 2020poster

We propose a differentiable sphere tracing algorithm to bridge the gap between inverse graphics methods and the recently proposed deep learning based implicit signed distance function. Due to the nature of the implicit function, the rendering process requires tremendous function queries, which is pa…

Cited by 350PDFcodeScholar
2020

Geometry-Aware Satellite-to-Ground Image Synthesis for Urban Areas

CVPR 2020poster

We present a novel method for generating panoramic street-view images which are geometrically consistent with a given satellite image. Different from existing approaches that completely rely on a deep learning architecture to generalize cross-view image distributions, our approach explicitly loops i…

Cited by 78PDFScholar
2020

OmniSLAM: Omnidirectional Localization and Dense Mapping for Wide-baseline Multi-camera Systems

ICRA 2020poster

In this paper, we present an omnidirectional localization and dense mapping system for a wide-baseline multiview stereo setup with ultra-wide field-of-view (FOV) fisheye cameras, which has a 360° coverage of stereo observations of the environment. For more practical and accurate reconstruction, we f…

Cited by 60SourceScholar
2020

Self-Supervised Human Depth Estimation From Monocular Videos

CVPR 2020poster

Previous methods on estimating detailed human depth often require supervised training with 'ground truth' depth data. This paper presents a self-supervised method that can be trained on YouTube videos without known depth, which makes training data collection simple and improves the generalization of…

Cited by 35PDFScholar
2019

DeepLiDAR: Deep Surface Normal Guided Depth Prediction for Outdoor Scene From Sparse LiDAR Data and Single Color Image

CVPR 2019poster

In this paper, we propose a deep learning architecture that produces accurate dense depth for the outdoor scene from a single color image and a sparse depth. Inspired by the indoor depth completion, our network estimates surface normals as the intermediate representation to produce dense depth, and…

Cited by 459PDFScholar
2019

Efficient 2D-3D Matching for Multi-Camera Visual Localization

ICRA 2019poster

Visual localization, i.e., determining the position and orientation of a vehicle with respect to a map, is a key problem in autonomous driving. We present a multi-camera visual inertial localization algorithm for large scale environments. To efficiently and effectively match features against a pre-b…

Cited by 41SourceScholar
2019

Project AutoVision: Localization and 3D Scene Perception for an Autonomous Vehicle with a Multi-Camera System

ICRA 2019poster

Project AutoVision aims to develop localization and 3D scene perception capabilities for a self-driving vehicle. Such capabilities will enable autonomous navigation in urban and rural environments, in day and night, and with cameras as the only exteroceptive sensors. The sensor suite employs many ca…

Cited by 149SourceScholar
2019

Real-Time Dense Mapping for Self-Driving Vehicles using Fisheye Cameras

ICRA 2019poster

We present a real-time dense geometric mapping algorithm for large-scale environments. Unlike existing methods which use pinhole cameras, our implementation is based on fisheye cameras whose large field of view benefits various computer vision applications for self-driving vehicles such as visual-in…

Cited by 47SourceScholar
2019

Reflection Separation using a Pair of Unpolarized and Polarized Images

NeurIPS 2019spotlight

When we take photos through glass windows or doors, the transmitted background scene is often blended with undesirable reflection. Separating two layers apart to enhance the image quality is of vital importance for both human and machine perception. In this paper, we propose to exploit physical cons…