← Search

Jiapeng Tang

14 accepted papers

2025

GAF: Gaussian Avatar Reconstruction from Monocular Videos via Multi-view Diffusion

CVPR 2025poster

We propose a novel approach for reconstructing animatable 3D Gaussian avatars from monocular videos captured by commodity devices like smartphones. Photorealistic 3D head avatar reconstruction from such recordings is challenging due to limited observations, which leaves unobserved regions under-cons…

2025

ROGR: Relightable 3D Objects using Generative Relighting

NeurIPS 2025spotlight

We introduce ROGR, a novel approach that reconstructs a relightable 3D model of an object captured from multiple views, driven by a generative relighting model that simulates the effects of placing the object under novel environment illuminations. Our method samples the appearance of the object unde…

Cited by 0SourceScholar
2025

SHeaP: Self-Supervised Head Geometry Predictor Learned via 2D Gaussians

ICCV 2025poster

Accurate, real-time 3D reconstruction of human heads from monocular images and videos underlies numerous visual applications. As 3D ground truth data is hard to come by at scale, previous methods have sought to learn from abundant 2D videos in a self-supervised manner. Typically, this involves the u…

Cited by 0SourcePDFScholar
2024

DPHMs: Diffusion Parametric Head Models for Depth-based Tracking

CVPR 2024poster

We introduce Diffusion Parametric Head Models (DPHMs) a generative model that enables robust volumetric head reconstruction and tracking from monocular depth sequences. While recent volumetric head models such as NPHMs can now excel in representing high-fidelity head geometries tracking and reconstr…

Cited by 6SourcePDFScholar
2024

DiffuScene: Denoising Diffusion Models for Generative Indoor Scene Synthesis

CVPR 2024poster

We present DiffuScene for indoor 3D scene synthesis based on a novel scene configuration denoising diffusion model. It generates 3D instance properties stored in an unordered object set and retrieves the most similar geometry for each object configuration which is characterized as a concatenation of…

Cited by 99SourcePDFScholar
2024

KMTalk: Speech-Driven 3D Facial Animation with Key Motion Embedding

ECCV 2024poster

"We present a novel approach for synthesizing 3D facial motions from audio sequences using key motion embeddings. Despite recent advancements in data-driven techniques, accurately mapping between audio signals and 3D facial meshes remains challenging. Direct regression of the entire sequence often l…

2024

Motion2VecSets: 4D Latent Vector Set Diffusion for Non-rigid Shape Reconstruction and Tracking

CVPR 2024poster

We introduce Motion2VecSets a 4D diffusion model for dynamic surface reconstruction from point cloud sequences. While existing state-of-the-art methods have demonstrated success in reconstructing non-rigid objects using neural field representations conventional feed-forward networks encounter challe…

Cited by 9SourcePDFScholar
2023

RGBD2: Generative Scene Synthesis via Incremental View Inpainting Using RGBD Diffusion Models

CVPR 2023poster

We address the challenge of recovering an underlying scene geometry and colors from a sparse set of RGBD view observations. In this work, we present a new solution termed RGBD2 that sequentially generates novel RGBD views along a camera trajectory, and the scene geometry is simply the fusion result…

Cited by 36SourcePDFScholar
2021

Dual Attention Guided Gaze Target Detection in the Wild

CVPR 2021poster

Gaze target detection aims to infer where each person in a scene is looking. Existing works focus on 2D gaze and 2D saliency, but fail to exploit 3D contexts. In this work, we propose a three-stage method to simulate the human gaze inference behavior in 3D space. In the first stage, we introduce a c…

Cited by 91PDFcodeScholar
2021

Learning Parallel Dense Correspondence From Spatio-Temporal Descriptors for Efficient and Robust 4D Reconstruction

CVPR 2021poster

This paper focuses on the task of 4D shape reconstruction from a sequence of point clouds. Despite the recent success achieved by extending deep implicit representations into 4D space, it is still a great challenge in two respects, i.e. how to design a flexible framework for learning robust spatio-t…

Cited by 33PDFcodeScholar
2021

SA-ConvONet: Sign-Agnostic Optimization of Convolutional Occupancy Networks

ICCV 2021poster

Surface reconstruction from point clouds is a fundamental problem in the computer vision and graphics community. Recent state-of-the-arts solve this problem by individually optimizing each local implicit field during inference. Without considering the geometric relationships between local fields, th…

Cited by 87PDFcodeScholar
2019

A Skeleton-Bridged Deep Learning Approach for Generating Meshes of Complex Topologies From Single RGB Images

CVPR 2019oral

This paper focuses on the challenging task of learning 3D object surface reconstructions from single RGB images. Existing methods achieve varying degrees of success by using different geometric representations. However, they all have their own drawbacks, and cannot well reconstruct those surfaces of…

Cited by 104PDFScholar
2019

Deep Mesh Reconstruction From Single RGB Images via Topology Modification Networks

ICCV 2019poster

Reconstructing the 3D mesh of a general object from a single image is now possible thanks to the latest advances of deep learning technologies. However, due to the nontrivial difficulty of generating a feasible mesh structure, the state-of-the-art approaches often simplify the problem by learning th…

Cited by 246PDFcodeScholar