← Search

Yan-Pei Cao

44 accepted papers

2026

Beyond Reassembly: Fractured Object Recovery with Missing Parts

CVPR 2026

We propose a novel learning-based task named fractured object recovery. Unlike the previous fractured object reassembly task that only aligns existing parts with overlaps, our task aims to recover the complete shape by not only reassembling irrelevant parts but also predicting missing parts. Our tas

Cited by 0SourceScholar
2026

FACE: A Face-based Autoregressive Representation for High-Fidelity and Efficient Mesh Generation

CVPR 2026

Autoregressive models for 3D mesh generation suffer from a fundamental limitation: they flatten meshes into long vertex-coordinate sequences. This results in prohibitive computational costs, hindering the efficient synthesis of high-fidelity geometry. We argue this bottleneck stems from operating at

Cited by 0SourceScholar
2026

GeoSAM2: Unleashing the Power of SAM2 for 3D Part Segmentation

CVPR 2026

We introduce GeoSAM2, a prompt-controllable framework for 3D part segmentation that casts the task as multi-view 2D mask prediction. Given a textureless object, we render normal and point maps from predefined viewpoints and accept simple 2D prompts--clicks or boxes--to guide part selection. These pr

Cited by 0SourceScholar
2026

HoloPart: Generative 3D Part Amodal Segmentation

ICLR 2026poster

3D part amodal segmentation--decomposing a 3D shape into complete, semantically meaningful parts, even when occluded--is a challenging but crucial task for 3D content creation and understanding. Existing 3D part segmentation methods only identify visible surface patches, limiting their utility. Insp…

Cited by 0SourceScholar
2026

Lafite: A Generative Latent Field for 3D Native Texturing

CVPR 2026

Generating high-fidelity, seamless textures directly on 3D surfaces, a process we term 3D-native texturing, is a fundamental open challenge, promising to overcome the limitations of traditional UV-based and multi-view projection methods. While promising, existing native approaches are bottlenecked b

Cited by 0SourceScholar
2026

Stereo World Model: Camera-Guided Stereo Video Generation

CVPR 2026

We present StereoWorld, a camera-conditioned stereo world model that jointly learns appearance and binocular geometry for end-to-end stereo video generation.Unlike monocular RGB or RGBD approaches, StereoWorld operates exclusively within the RGB modality, while simultaneously grounding geometry dire

Cited by 0SourcecodeScholar
2026

Unified Camera Positional Encoding for Controlled Video Generation

CVPR 2026

Transformers have emerged as a universal backbone across 3D perception, video generation, and world models for autonomous driving and embodied AI, where understanding camera geometry is essential for grounding visual observations in three-dimensional space. However, existing camera encoding methods

Cited by 0SourcecodeScholar
2025

DI-PCG: Diffusion-based Efficient Inverse Procedural Content Generation for High-quality 3D Asset Creation

CVPR 2025poster

Procedural Content Generation (PCG) is powerful in creating high-quality 3D contents, yet controlling it to produce desired shapes is difficult and often requires extensive parameter tuning. Inverse Procedural Content Generation aims to automatically find the best parameters under the input conditio…

Cited by 4SourcePDFScholar
2025

Deformable Radial Kernel Splatting

CVPR 2025poster

Recently, Gaussian splatting has emerged as a robust technique for representing 3D scenes, enabling real-time rasterization and high-fidelity rendering. However, Gaussians' inherent radial symmetry and smoothness constraints limit their ability to represent complex shapes, often requiring thousands…

Cited by 1SourcePDFScholar
2025

GCRayDiffusion: Pose-Free Surface Reconstruction via Geometric Consistent Ray Diffusion

ICCV 2025poster

Accurate surface reconstruction from unposed images is crucial for efficient 3D object or scene creation. However, it remains challenging, particularly for the joint camera pose estimation. Previous approaches have achieved impressive pose-free surface reconstruction results in dense-view settings,…

2025

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

CVPR 2025poster

This paper introduces MIDI, a novel paradigm for compositional 3D scene generation from a single image. Unlike existing methods that rely on reconstruction or retrieval techniques or recent approaches that employ multi-stage object-by-object generation, MIDI extends pre-trained image-to-3D object ge…

Cited by 1SourcePDFScholar
2025

MV-Adapter: Multi-View Consistent Image Generation Made Easy

ICCV 2025poster

Existing multi-view image generation methods often make invasive modifications to pre-trained text-to-image (T2I) models and require full fine-tuning, leading to high computational costs and degradation in image quality due to scarce high-quality 3D data. This paper introduces MV-Adapter, an efficie…

Cited by 0SourcePDFScholar
2025

NeuFrameQ: Neural Frame Fields for Scalable and Generalizable Anisotropic Quadrangulation

ICCV 2025poster

Quad meshes play a crucial role in computer graphics applications, yet automatically generating high-quality quad meshes remains challenging. Traditional quadrangulation approaches rely on local geometric features and manual constraints, often producing suboptimal mesh layouts that fail to capture g…

Cited by 0SourcePDFScholar
2025

PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit Remeshing

CVPR 2025poster

Photorealistic 3D human modeling is essential for various applications and has seen tremendous progress. However, existing methods for monocular full-body reconstruction, typically relying on front and/or predicted back view, still struggle with satisfactory performance due to the ill-posed nature o…

2025

SparseFlex: High-Resolution and Arbitrary-Topology 3D Shape Modeling

ICCV 2025poster

Creating high-fidelity 3D meshes with arbitrary topology, including open surfaces and complex interiors, remains a significant challenge. Existing implicit field methods often require costly and detail-degrading watertight conversion, while other approaches struggle with high resolutions. This paper…

2025

SuperMat: Physically Consistent PBR Material Estimation at Interactive Rates

ICCV 2025poster

Decomposing physically-based materials from images into their constituent properties remains challenging, particularly when maintaining both computational efficiency and physical consistency. While recent diffusion-based approaches have shown promise, they face substantial computational overhead due…

2024

DMiT: Deformable Mipmapped Tri-Plane Representation for Dynamic Scenes

ECCV 2024poster

"Neural Radiance Fields (NeRF) have achieved remarkable progress on dynamic scenes with deformable objects. Nonetheless, most previous works required multi-view inputs or long training time (several hours), making it hard to apply them for real-world scenarios. Recent works dedicated to addressing b…

Cited by 0SourcePDFScholar
2024

DreamAvatar: Text-and-Shape Guided 3D Human Avatar Generation via Diffusion Models

CVPR 2024poster

We present DreamAvatar a text-and-shape guided framework for generating high-quality 3D human avatars with controllable poses. While encouraging results have been reported by recent methods on text-guided 3D common object generation generating high-quality human avatars remains an open challenge due…

2024

DreamDiffusion: High-Quality EEG-to-Image Generation with Temporal Masked Signal Modeling and CLIP Alignment

ECCV 2024poster

"This paper introduces DreamDiffusion, a novel method for generating high-quality images directly from brain electroencephalogram (EEG) signals, without the need to translate thoughts into text. DreamDiffusion leverages pre-trained text-to-image models and employs temporal masked signal modeling to…

2024

DynVideo-E: Harnessing Dynamic NeRF for Large-Scale Motion- and View-Change Human-Centric Video Editing

CVPR 2024poster

Despite recent progress in diffusion-based video editing existing methods are limited to short-length videos due to the contradiction between long-range consistency and frame-wise editing. Prior attempts to address this challenge by introducing video-2D representations encounter significant difficul…

2024

EpiDiff: Enhancing Multi-View Synthesis via Localized Epipolar-Constrained Diffusion

CVPR 2024poster

Generating multiview images from a single view facilitates the rapid generation of a 3D mesh conditioned on a single image. Recent methods that introduce 3D global representation into diffusion models have shown the potential to generate consistent multiviews but they have reduced generation speed a…

2024

HiFi-123: Towards High-fidelity One Image to 3D Content Generation

ECCV 2024poster

"Recent advances in diffusion models have enabled 3D generation from a single image. However, current methods often produce suboptimal results for novel views, with blurred textures and deviations from the reference image, limiting their practical applications. In this paper, we introduce HiFi-123,…

Cited by 26SourcePDFScholar
2024

SC-GS: Sparse-Controlled Gaussian Splatting for Editable Dynamic Scenes

CVPR 2024poster

Novel view synthesis for dynamic scenes is still a challenging problem in computer vision and graphics. Recently Gaussian splatting has emerged as a robust technique to represent static scenes and enable high-quality and real-time novel view synthesis. Building upon this technique we propose a new r…

2024

SC-NeuS: Consistent Neural Surface Reconstruction from Sparse and Noisy Views

AAAI 2024technical

The recent neural surface reconstruction approaches using volume rendering have made much progress by achieving impressive surface reconstruction quality, but are still limited to dense and highly accurate posed views. To overcome such drawbacks, this paper pays special attention on the consistent s…

2024

Sparse3D: Distilling Multiview-Consistent Diffusion for Object Reconstruction from Sparse Views

AAAI 2024technical

Reconstructing 3D objects from extremely sparse views is a long-standing and challenging problem. While recent techniques employ image diffusion models for generating plausible images at novel viewpoints or for distilling pre-trained diffusion priors into 3D representations using score distillation…

Cited by 26SourcePDFScholar
2024

SparseGNV: Generating Novel Views of Indoor Scenes with Sparse RGB-D Images

AAAI 2024technical

We study to generate novel views of indoor scenes given sparse input views. The challenge is to achieve both photorealism and view consistency. We present SparseGNV: a learning framework that incorporates 3D structures and image generative models to generate novel views with three modules. The first…

Cited by 0SourcePDFScholar
2024

Splatter a Video: Video Gaussian Representation for Versatile Processing

NeurIPS 2024poster

Video representation is a long-standing problem that is crucial for various downstream tasks, such as tracking, depth prediction, segmentation, view synthesis, and editing. However, current methods either struggle to model complex motions due to the absence of 3D structure or rely on implicit 3D rep…

Cited by 7SourcePDFScholar
2024

Triplane Meets Gaussian Splatting: Fast and Generalizable Single-View 3D Reconstruction with Transformers

CVPR 2024poster

Recent advancements in 3D reconstruction from single images have been driven by the evolution of generative models. Prominent among these are methods based on Score Distillation Sampling (SDS) and the adaptation of diffusion models in the 3D domain. Despite their progress these techniques often face…

2024

UniDream: Unifying Diffusion Priors for Relightable Text-to-3D Generation

ECCV 2024poster

"Recent advancements in text-to-3D generation technology have significantly advanced the conversion of textual descriptions into imaginative well-geometrical and finely textured 3D objects. Despite these developments, a prevalent limitation arises from the use of RGB data in diffusion or reconstruct…

2023

CL-NeRF: Continual Learning of Neural Radiance Fields for Evolving Scene Representation

NeurIPS 2023poster

Existing methods for adapting Neural Radiance Fields (NeRFs) to scene changes require extensive data capture and model retraining, which is both time-consuming and labor-intensive. In this paper, we tackle the challenge of efficiently adapting NeRFs to real-world scene changes over time using a few…

Cited by 10SourcePDFScholar
2023

Dream3D: Zero-Shot Text-to-3D Synthesis Using 3D Shape Prior and Text-to-Image Diffusion Models

CVPR 2023poster

Recent CLIP-guided 3D optimization methods, such as DreamFields and PureCLIPNeRF, have achieved impressive results in zero-shot text-to-3D synthesis. However, due to scratch training and random initialization without prior knowledge, these methods often fail to generate accurate and faithful 3D stru…

2023

HOSNeRF: Dynamic Human-Object-Scene Neural Radiance Fields from a Single Video

ICCV 2023poster

We introduce HOSNeRF, a novel 360deg free-viewpoint rendering method that reconstructs neural radiance fields for dynamic human-object-scene from a single monocular in-the-wild video. Our method enables pausing the video at any frame and rendering all scene details (dynamic humans, objects, and back…

Cited by 31PDFcodeScholar
2023

HRDFuse: Monocular 360deg Depth Estimation by Collaboratively Learning Holistic-With-Regional Depth Distributions

CVPR 2023poster

Depth estimation from a monocular 360 image is a burgeoning problem owing to its holistic sensing of a scene. Recently, some methods, e.g., OmniFusion, have applied the tangent projection (TP) to represent a 360 image and predicted depth values via patch-wise regressions, which are merged to get a d…

Cited by 32SourcePDFScholar
2023

OmniZoomer: Learning to Move and Zoom in on Sphere at High-Resolution

ICCV 2023poster

Omnidirectional images (ODIs) have become increasingly popular, as their large field-of-view (FoV) can offer viewers the chance to freely choose the view directions in immersive environments such as virtual reality. The Mobius transformation is typically employed to further provide the opportunity f…

Cited by 10PDFcodeScholar
2023

PanoGRF: Generalizable Spherical Radiance Fields for Wide-baseline Panoramas

NeurIPS 2023poster

Achieving an immersive experience enabling users to explore virtual environments with six degrees of freedom (6DoF) is essential for various applications such as virtual reality (VR). Wide-baseline panoramas are commonly used in these applications to reduce network bandwidth and storage requirements…

Cited by 9SourcePDFScholar
2023

Speech2Lip: High-fidelity Speech to Lip Generation by Learning from a Short Video

ICCV 2023poster

Synthesizing realistic videos according to a given speech is still an open challenge. Previous works have been plagued by issues such as inaccurate lip shape generation and poor image quality. The key reason is that only motions and appearances on limited facial areas (e.g., lip area) are mainly dri…

Cited by 17PDFcodeScholar
2023

SurfelNeRF: Neural Surfel Radiance Fields for Online Photorealistic Reconstruction of Indoor Scenes

CVPR 2023poster

Online reconstructing and rendering of large-scale indoor scenes is a long-standing challenge. SLAM-based methods can reconstruct 3D scene geometry progressively in real time but can not render photorealistic results. While NeRF-based methods produce promising novel view synthesis results, their lon…

Cited by 39SourcePDFScholar
2022

DeVRF: Fast Deformable Voxel Radiance Fields for Dynamic Scenes

NeurIPS 2022accept

Modeling dynamic scenes is important for many applications such as virtual reality and telepresence. Despite achieving unprecedented fidelity for novel view synthesis in dynamic scenes, existing methods based on Neural Radiance Fields (NeRF) suffer from slow convergence (i.e., model training time me…

2022

DoubleField: Bridging the Neural Surface and Radiance Fields for High-Fidelity Human Reconstruction and Rendering

CVPR 2022poster

We introduce DoubleField, a novel framework combining the merits of both surface field and radiance field for high-fidelity human reconstruction and rendering. Within DoubleField, the surface field and radiance field are associated together by a shared feature embedding and a surface-guided sampling…

Cited by 186PDFScholar
2021

Cycle4Completion: Unpaired Point Cloud Completion Using Cycle Transformation With Missing Region Coding

CVPR 2021poster

In this paper, we present a novel unpaired point cloud completion network, named Cycle4Completion, to infer the complete geometries from a partial 3D object. Previous unpaired completion methods merely focus on the learning of geometric correspondence from incomplete shapes to complete shapes, and i…

Cited by 134PDFcodeScholar
2021

PMP-Net: Point Cloud Completion by Learning Multi-Step Point Moving Paths

CVPR 2021poster

The task of point cloud completion aims to predict the missing part for an incomplete 3D shape. A widely used strategy is to generate a complete point cloud from the incomplete one. However, the unordered nature of point clouds will degrade the generation of high-quality 3D shapes, as the detailed t…

Cited by 235PDFcodeScholar
2021

SnowflakeNet: Point Cloud Completion by Snowflake Point Deconvolution With Skip-Transformer

ICCV 2021poster

Point cloud completion aims to predict a complete shape in high accuracy from its partial observation. However, previous methods usually suffered from discrete nature of point cloud and unstructured prediction of points in local regions, which makes it hard to reveal fine local geometric details on…

Cited by 326PDFcodeScholar
2019

Probabilistic Projective Association and Semantic Guided Relocalization for Dense Reconstruction

ICRA 2019poster

We present a real-time dense mapping system which uses the predicted 2D semantic labels for optimizing the geometric quality of reconstruction. With a combination of Convolutional Neural Networks (CNNs) for 2D labeling and a Simultaneous Localization and Mapping (SLAM) system for camera trajectory e…

Cited by 14SourceScholar
2018

Learning to Reconstruct High-quality 3D Shapes with Cascaded Fully Convolutional Networks

ECCV 2018poster

We present a data-driven approach to reconstructing high-resolution and detailed volumetric representations of 3D shapes. Although well studied, algorithms for volumetric fusion from multi-view depth scans are still prone to scanning noise and occlusions, making it hard to obtain high-fidelity 3D re…

Cited by 38SourcePDFScholar