← Search

Zeng Huang

14 accepted papers

2026

Talking Together: Synthesizing Co-Located 3D Conversations from Audio

CVPR 2026

We tackle the challenging task of generating complete 3D facial animations for two interacting, co-located participants from a mixed audio stream. While existing methods often produce disembodied "talking heads" akin to a video conference call, our work is the first to explicitly model the dynamic 3

Cited by 0SourceScholar
2024

Loc3Diff: Local Diffusion for 3D Human Head Synthesis and Editing

ECCV 2024poster

"We present a novel framework for generating photorealistic 3D human head and subsequently manipulating and reposing them with remarkable flexibility. The proposed approach constructs an implicit representation of 3D human heads, anchored on a parametric face model. To enhance representational capab…

Cited by 0SourcePDFScholar
2023

Learning Personalized High Quality Volumetric Head Avatars From Monocular RGB Videos

CVPR 2023poster

We propose a method to learn a high-quality implicit 3D head avatar from a monocular RGB video captured in the wild. The learnt avatar is driven by a parametric face model to achieve user-controlled facial expressions and head poses. Our hybrid pipeline combines the geometry prior and dynamic tracki…

Cited by 20SourcePDFScholar
2022

Cross-Modal 3D Shape Generation and Manipulation

ECCV 2022poster

"Creating and editing the shape and color of 3D objects require tremendous human effort and expertise. Compared to direct manipulation in 3D interfaces, 2D interactions such as sketches and scribbles are usually much more natural and intuitive for the users. In this paper, we propose a generic multi…

Cited by 34SourcePDFScholar
2022

R2L: Distilling Neural Radiance Field to Neural Light Field for Efficient Novel View Synthesis

ECCV 2022poster

"Recent research explosion on Neural Radiance Field (NeRF) shows the encouraging potential to represent complex scenes with neural networks. One major drawback of NeRF is its prohibitive inference time: Rendering a single pixel requires querying the NeRF network hundreds of times. To resolve it, exi…

2021

S3: Neural Shape, Skeleton, and Skinning Fields for 3D Human Modeling

CVPR 2021poster

Constructing and animating humans is an important component for building virtual worlds in a wide variety of applications such as virtual reality or robotics testing in simulation. As there are exponentially many variations of humans with different shape, pose and clothing, it is critical to develop…

Cited by 85PDFScholar
2020

Monocular Real-Time Volumetric Performance Capture

ECCV 2020poster

We present the first approach to volumetric performance capture and novel-view rendering at real-time speed from monocular video, eliminating the need for expensive multi-view systems or cumbersome pre-acquisition of a personalized template model. Our system reconstructs a fully textured 3D human fr…

2019

Learning Perspective Undistortion of Portraits

ICCV 2019oral

Near-range portrait photographs often contain perspective distortion artifacts that bias human perception and challenge both facial recognition and reconstruction techniques. We present the first deep learning based approach to remove such artifacts from unconstrained portraits. In contrast to the p…

Cited by 33PDFScholar
2019

PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization

ICCV 2019poster

We introduce Pixel-aligned Implicit Function (PIFu), an implicit representation that locally aligns pixels of 2D images with the global context of their corresponding 3D object. Using PIFu, we propose an end-to-end deep learning method for digitizing highly detailed clothed humans that can infer bot…

Cited by 1438PDFScholar
2018

Auto-Conditioned Recurrent Networks for Extended Complex Human Motion Synthesis

ICLR 2018poster

We present a real-time method for synthesizing highly complex human motions using a novel training regime we call the auto-conditioned Recurrent Neural Network (acRNN). Recently, researchers have attempted to synthesize new motion by using autoregressive techniques, but existing methods tend to free…

Cited by 264SourcePDFScholar
2018

Deep Volumetric Video From Very Sparse Multi-View Performance Capture

ECCV 2018poster

We present a deep learning-based volumetric capture approach for performance capture using a passive and highly sparse multi-view capture system. We focus on a template-free, per-frame 3D surface reconstruction from as few as three RGB sensors, where conventional visual hull or multi-view stereo met…

Cited by 144SourcePDFScholar
2017

Realistic Dynamic Facial Textures From a Single Image Using GANs

ICCV 2017poster

We present a novel method to realistically puppeteer and animate a face from a single RGB image using a source video sequence. We begin by fitting a multilinear PCA model to obtain the 3D geometry and a single texture of the target face. In order for the animation to be realistic, however, we need d…

Cited by 116PDFScholar