← Search

Hongyi Xu

15 accepted papers

2025

DiffPortrait360: Consistent Portrait Diffusion for 360 View Synthesis

CVPR 2025poster

Generating high-quality 360-degree views of human heads from single-view images is essential for enabling accessible immersive telepresence applications and scalable personalized content creation.While cutting-edge methods for full head generation are limited to modeling realistic human heads, the l…

2025

X-Dancer: Expressive Music to Human Dance Video Generation

ICCV 2025poster

We present X-Dancer, a novel zero-shot music-driven image animation pipeline that creates diverse and long-range lifelike human dance videos from a single static image. As its core, we introduce a unified transformer-diffusion framework, featuring an autoregressive transformer model that synthesize…

Cited by 0SourcePDFScholar
2025

X-Dyna: Expressive Dynamic Human Image Animation

CVPR 2025highlight

We introduce X-Dyna, a novel zero-shot, diffusion-based pipeline for animating a single human image using facial expressions and body movements derived from a driving video, that generates realistic, context-aware dynamics for both the subject and the surrounding environment. Building on prior appro…

2025

X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention

ICLR 2025poster

We propose X-NeMo, a novel zero-shot diffusion-based portrait animation pipeline that animates a static portrait using facial movements from a driving video of a different individual. Our work first identifies the root causes of the limitations in prior approaches, such as identity leakage and diffi…

Cited by 0SourcePDFScholar
2024

DiffPortrait3D: Controllable Diffusion for Zero-Shot Portrait View Synthesis

CVPR 2024highlight

We present DiffPortrait3D a conditional diffusion model that is capable of synthesizing 3D-consistent photo-realistic novel views from as few as a single in-the-wild portrait. Specifically given a single RGB input we aim to synthesize plausible but consistent facial details rendered from novel camer…

2024

MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion

ICML 2024poster

In this work, we propose MagicPose, a diffusion-based model for 2D human pose and facial expression retargeting. Specifically, given a reference image, we aim to generate a person's new images by controlling the poses and facial expressions while keeping the identity unchanged. To this end, we propo…

2023

GETAvatar: Generative Textured Meshes for Animatable Human Avatars

ICCV 2023poster

We study the problem of 3D-aware full-body human generation, aiming at creating animatable human avatars with high-quality textures and geometries. Generally, two challenges remain in this field: i) existing methods struggle to generate geometries with rich realistic details such as the wrinkles of…

Cited by 23PDFScholar
2023

OmniAvatar: Geometry-Guided Controllable 3D Head Synthesis

CVPR 2023poster

We present OmniAvatar, a novel geometry-guided 3D head synthesis model trained from in-the-wild unstructured images that is capable of synthesizing diverse identity-preserved 3D heads with compelling dynamic details under full disentangled control over camera poses, facial expressions, head shapes,…

Cited by 29SourcePDFScholar
2023

PanoHead: Geometry-Aware 3D Full-Head Synthesis in 360deg

CVPR 2023poster

Synthesis and reconstruction of 3D human head has gained increasing interests in computer vision and computer graphics recently. Existing state-of-the-art 3D generative adversarial networks (GANs) for 3D human head synthesis are either limited to near-frontal views or hard to preserve 3D consistency…

2022

Feature Generation and Hypothesis Verification for Reliable Face Anti-spoofing

AAAI 2022technical

Although existing face anti-spoofing (FAS) methods achieve high accuracy in intra-domain experiments, their effects drop severely in cross-domain scenarios because of poor generalization. Recently, multifarious techniques have been explored, such as domain generalization and representation disentang…

2022

Trajectory Optimization for Physics-Based Reconstruction of 3D Human Pose From Monocular Video

CVPR 2022poster

We focus on the task of estimating a physically plausible articulated human motion from monocular video. Existing approaches that do not consider physics often produce temporally inconsistent output with motion artifacts, while state-of-the-art physics-based approaches have either been shown to work…

Cited by 48PDFScholar
2021

H-NeRF: Neural Radiance Fields for Rendering and Temporal Reconstruction of Humans in Motion

NeurIPS 2021spotlight

We present neural radiance fields for rendering and temporal (4D) reconstruction of humans in motion (H-NeRF), as captured by a sparse set of cameras or even from a monocular video. Our approach combines ideas from neural scene representation, novel-view synthesis, and implicit statistical geometric…

Cited by 205SourcePDFScholar
2021

imGHUM: Implicit Generative Models of 3D Human Shape and Articulated Pose

ICCV 2021poster

We present imGHUM, the first holistic generative model of 3D human shape and articulated pose, represented as a signed distance function. In contrast to prior work, we model the full human body implicitly as a function zero-level-set and without the use of an explicit template mesh. We propose a nov…

Cited by 121PDFcodeScholar
2020

GHUM & GHUML: Generative 3D Human Shape and Articulated Pose Models

CVPR 2020oral

We present a statistical, articulated 3D human shape modeling pipeline, within a fully trainable, modular, deep learning framework. Given high-resolution complete 3D body scans of humans, captured in various poses, together with additional closeups of their head and facial expressions, as well as ha…

Cited by 423PDFcodeScholar
2020

Weakly Supervised 3D Human Pose and Shape Reconstruction with Normalizing Flows

ECCV 2020poster

Monocular 3D human pose and shape estimation is challenging due to the many degrees of freedom of the human body and the difficulty to acquire training data for large-scale supervised learning in complex visual scenes where humans with diverse shape and appearance, appear against complex backgrounds…

Cited by 160SourcePDFScholar