← Search

Zerong Zheng

29 accepted papers

2026

Instilling an Active Mind in Avatars via Cognitive Simulation

ICLR 2026oral

Current video avatar models can generate fluid animations but struggle to capture a character's authentic essence, primarily synchronizing motion with low-level audio cues instead of understanding higher-level semantics like emotion or intent. To bridge this gap, we propose a novel framework for gen…

Cited by 0SourcecodeScholar
2026

InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions

ICLR 2026poster

End-to-end human animation with rich multi-modal conditions, e.g., text, image and audio has achieved remarkable advancements in recent years. However, most existing methods could only animate a single subject and inject conditions in a global manner, ignoring scenarios that multiple concepts could…

Cited by 0SourceScholar
2025

CyberHost: A One-stage Diffusion Framework for Audio-driven Talking Body Generation

ICLR 2025oral

Diffusion-based video generation technology has advanced significantly, catalyzing a proliferation of research in human animation. While breakthroughs have been made in driving human animation through various modalities for portraits, most of current solutions for human body animation still focus on…

Cited by 0SourcePDFScholar
2025

OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

ICCV 2025poster

End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generation models, limiting their potential in real applications. In this paper, we propo…

Cited by 0SourcePDFScholar
2024

Animatable Gaussians: Learning Pose-dependent Gaussian Maps for High-fidelity Human Avatar Modeling

CVPR 2024poster

Modeling animatable human avatars from RGB videos is a long-standing and challenging problem. Recent works usually adopt MLP-based neural radiance fields (NeRF) to represent 3D humans but it remains difficult for pure MLPs to regress pose-dependent garment details. To this end we introduce Animatabl…

2024

Control4D: Efficient 4D Portrait Editing with Text

CVPR 2024poster

We introduce Control4D an innovative framework for editing dynamic 4D portraits using text instructions. Our method addresses the prevalent challenges in 4D editing notably the inefficiencies of existing 4D representations and the inconsistent editing effect caused by diffusion-based editors. We fir…

Cited by 22SourcePDFScholar
2024

DiffPerformer: Iterative Learning of Consistent Latent Guidance for Diffusion-based Human Video Generation

CVPR 2024poster

Existing diffusion models for pose-guided human video generation mostly suffer from temporal inconsistency in the generated appearance and poses due to the inherent randomization nature of the generation process. In this paper we propose a novel framework DiffPerformer to synthesize high-fidelity an…

Cited by 1SourcePDFScholar
2024

Gaussian Head Avatar: Ultra High-fidelity Head Avatar via Dynamic Gaussians

CVPR 2024poster

Creating high-fidelity 3D head avatars has always been a research hotspot but there remains a great challenge under lightweight sparse view setups. In this paper we propose Gaussian Head Avatar represented by controllable 3D Gaussians for high-fidelity head avatar modeling. We optimize the neutral 3…

2024

MeshAvatar: Learning High-quality Triangular Human Avatars from Multi-view Videos

ECCV 2024poster

"We present a novel pipeline for learning high-quality triangular human avatars from multi-view videos. Recent methods for avatar learning are typically based on neural radiance fields (NeRF), which is not compatible with traditional graphics pipeline and poses great challenges for operations like e…

2024

RAM-Avatar: Real-time Photo-Realistic Avatar from Monocular Videos with Full-body Control

CVPR 2024poster

This paper focuses on advancing the applicability of human avatar learning methods by proposing RAM-Avatar which learns a Real-time photo-realistic Avatar that supports full-body control from Monocular videos. To achieve this goal RAM-Avatar leverages two statistical templates responsible for modeli…

Cited by 3SourcePDFScholar
2023

CloSET: Modeling Clothed Humans on Continuous Surface With Explicit Template Decomposition

CVPR 2023poster

Creating animatable avatars from static scans requires the modeling of clothing deformations in different poses. Existing learning-based methods typically add pose-dependent deformations upon a minimally-clothed mesh template or a learned implicit template, which have limitations in capturing detail…

Cited by 28SourcePDFScholar
2023

Leveraging Intrinsic Properties for Non-Rigid Garment Alignment

ICCV 2023poster

We address the problem of aligning real-world 3D data of garments, which benefits many applications such as texture learning, physical parameter estimation, generative modeling of garments, etc. Existing extrinsic methods typically perform non-rigid iterative closest point and struggle to align deta…

Cited by 7PDFcodeScholar
2023

Tensor4D: Efficient Neural 4D Decomposition for High-Fidelity Dynamic Reconstruction and Rendering

CVPR 2023highlight

We present Tensor4D, an efficient yet effective approach to dynamic scene modeling. The key of our solution is an efficient 4D tensor decomposition method so that the dynamic scene can be directly represented as a 4D spatio-temporal tensor. To tackle the accompanying memory issue, we decompose the 4…

2022

AvatarCap: Animatable Avatar Conditioned Monocular Human Volumetric Capture

ECCV 2022poster

"To address the ill-posed problem caused by partial observations in monocular human volumetric capture, we present AvatarCap, a novel framework that introduces animatable avatars into the capture pipeline for high-fidelity reconstruction in both visible and invisible regions. Our method firstly crea…

2022

DiffuStereo: High Quality Human Reconstruction via Diffusion-Based Stereo Using Sparse Cameras

ECCV 2022poster

"We propose DiffuStereo, a novel system using only sparse cameras (8 in this work) for high-quality 3D human reconstruction. At its core is a novel diffusion-based stereo module, which introduces diffusion models, a type of powerful generative models, into the iterative stereo matching network. To t…

Cited by 68SourcePDFScholar
2022

High-Fidelity Human Avatars From a Single RGB Camera

CVPR 2022poster

In this paper, we propose a coarse-to-fine framework to reconstruct a personalized high-fidelity human avatar from a monocular video. To deal with the misalignment problem caused by the changed poses and shapes in different frames, we design a dynamic surface network to recover pose-dependent surfac…

Cited by 40PDFScholar
2022

Learning Implicit Templates for Point-Based Clothed Human Modeling

ECCV 2022poster

"We present FITE, a First-Implicit-Then-Explicit framework for modeling human avatars in clothing. Our framework first learns implicit surface templates representing the coarse clothing topology, and then employs the templates to guide the generation of point sets which further capture pose-dependen…

2022

Structured Local Radiance Fields for Human Avatar Modeling

CVPR 2022poster

It is extremely challenging to create an animatable clothed human avatar from RGB videos, especially for loose clothes due to the difficulties in motion modeling. To address this problem, we introduce a novel representation on the basis of recent neural scene rendering techniques. The core of our re…

Cited by 218PDFcodeScholar
2021

DeepMultiCap: Performance Capture of Multiple Characters Using Sparse Multiview Cameras

ICCV 2021poster

We propose DeepMultiCap, a novel method for multi-person performance capture using sparse multi-view cameras. Our method can capture time varying surface details without the need of using pre-scanned template models. To tackle with the serious occlusion challenge for close interacting scenes, we com…

Cited by 110PDFScholar
2021

Function4D: Real-Time Human Volumetric Capture From Very Sparse Consumer RGBD Sensors

CVPR 2021poster

Human volumetric capture is a long-standing topic in computer vision and computer graphics. Although high-quality results can be achieved using sophisticated off-line systems, real-time human volumetric capture of complex scenarios, especially using light-weight setups, remains challenging. In this…

Cited by 361PDFScholar
2021

POSEFusion: Pose-Guided Selective Fusion for Single-View Human Volumetric Capture

CVPR 2021poster

We propose POse-guided SElective Fusion (POSEFusion), a single-view human volumetric capture method that leverages tracking-based methods and tracking-free inference to achieve high-fidelity and dynamic 3D reconstruction. By contributing a novel reconstruction framework which contains pose-guided ke…

Cited by 33PDFScholar
2020

RobustFusion: Human Volumetric Capture with Data-driven Visual Cues using a RGBD Camera

ECCV 2020poster

High-quality and complete 4D reconstruction of human activities is critical for immersive VR/AR experience, but it suffers from inherent self-scanning constraint and consequent fragile tracking under the monocular setting. In this paper, inspired by the huge potential of learning-based human modelin…

Cited by 106SourcePDFScholar
2019

SimulCap : Single-View Human Performance Capture With Cloth Simulation

CVPR 2019poster

This paper proposes a new method for live free-viewpoint human performance capture with dynamic details (e.g., cloth wrinkles) using a single RGBD camera. Our main contributions are: (i) a multi-layer representation of garments and body, and (ii) a physics-based performance capture procedure. We fir…

Cited by 125PDFScholar
2018

DoubleFusion: Real-Time Capture of Human Performances With Inner Body Shapes From a Single Depth Sensor

CVPR 2018poster

We propose DoubleFusion, a new real-time system that combines volumetric dynamic reconstruction with data-driven template fitting to simultaneously reconstruct detailed geometry, non-rigid motion and the inner human body shape from a single depth camera. One of the key contributions of this method i…

Cited by 371SourcePDFScholar
2018

HybridFusion: Real-Time Performance Capture Using a Single Depth Sensor and Sparse IMUs

ECCV 2018poster

We propose a light-weight and highly robust real-time human performance capture method based on a single depth camera and sparse inertial measurement units (IMUs). The proposed method combines non-rigid surface tracking and volumetric surface fusion to simultaneously reconstruct challenging motions,…

Cited by 112SourcePDFScholar