← Search

Hongwen Zhang

36 accepted papers

2026

GAF: Gaussian Action Field As a 4D Representation for Dynamic World Modeling in Robotic Manipulation

ICRA 2026poster

Accurate scene perception is critical for vision-based robotic manipulation. Existing approaches typically follow either a Vision-to-Action V-A paradigm, predicting actions directly from visual inputs, or a Vision-to-3D-to-Action V-3D-A paradigm, leveraging intermediate 3D representations. However, …

2026

HOSIG: Full-Body Human-Object-Scene Interaction Generation with Hierarchical Scene Perception

AAAI 2026technical

Generating high-fidelity full-body human interactions with dynamic objects and static scenes remains a critical challenge in computer graphics and animation. Existing methods for human-object interaction often neglect scene context, leading to implausible penetrations, while human-scene interaction

Cited by 0SourcePDFScholar
2026

MetricHMSR: Metric Human Mesh and Scene Recovery from Monocular Images

CVPR 2026

We introduce MetricHMSR (Metric Human Mesh and Scene Recovery), a novel approach for metric human mesh and scene recovery from monocular images. Due to unrealistic assumptions in the camera model and inherent challenges in metric perception, existing approaches struggle to achieve human pose and met

Cited by 0SourcecodeScholar
2026

RECOM: REALISTIC CO-SPEECH MOTION GENERATION WITH RECURRENT EMBEDDED TRANSFORMER

ICASSP 2026poster

We present ReCoM, an efficient framework for generating high-fidelity and generalizable human body motions synchronized with speech. The core innovation lies in the Recurrent Embedded Transformer (RET), which integrates Dynamic Embedding Regularization (DER) into a Vision Transformer (ViT) core arch…

Cited by 0SourcePDFScholar
2026

SharpTimeGS: Sharp and Stable Dynamic Gaussian Splatting via Lifespan Modulation

CVPR 2026

Novel view synthesis of dynamic scenes is fundamental to achieving photorealistic 4D reconstruction and immersive visual experiences. Recent progress in Gaussian-based representations has significantly improved real-time rendering quality, yet existing methods still struggle to maintain a balance be

Cited by 0SourceScholar
2025

HADES: Human Avatar with Dynamic Explicit Hair Strands

ICCV 2025poster

We introduce HADES, the first framework to seamlessly integrate dynamic hair into human avatars. HADES represents hair as strands bound to 3D Gaussians, with roots attached to the scalp. By modeling inertial and velocity-aware motion, HADES is able to simulate realistic hair dynamics that naturally…

Cited by 0SourcePDFScholar
2025

ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping

CVPR 2025highlight

In this paper, we introduce ManiVideo, a novel method for generating consistent and temporally coherent bimanual hand-object manipulation videos from given motion sequences of hands and objects. The core idea of ManiVideo is the construction of a multi-layer occlusion (MLO) representation that learn…

Cited by 1SourcePDFScholar
2025

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios

NeurIPS 2025spotlight

Hand-Object Interaction (HOI) generation has significant application potential. However, current 3D HOI motion generation approaches heavily rely on predefined 3D object models and lab-captured motion data, limiting generalization capabilities. Meanwhile, HOI video generation methods prioritize pixe…

Cited by 0SourcecodeScholar
2024

Control4D: Efficient 4D Portrait Editing with Text

CVPR 2024poster

We introduce Control4D an innovative framework for editing dynamic 4D portraits using text instructions. Our method addresses the prevalent challenges in 4D editing notably the inefficiencies of existing 4D representations and the inconsistent editing effect caused by diffusion-based editors. We fir…

Cited by 22SourcePDFScholar
2024

Gaussian Head Avatar: Ultra High-fidelity Head Avatar via Dynamic Gaussians

CVPR 2024poster

Creating high-fidelity 3D head avatars has always been a research hotspot but there remains a great challenge under lightweight sparse view setups. In this paper we propose Gaussian Head Avatar represented by controllable 3D Gaussians for high-fidelity head avatar modeling. We optimize the neutral 3…

2024

GaussianAvatar: Towards Realistic Human Avatar Modeling from a Single Video via Animatable 3D Gaussians

CVPR 2024poster

We present GaussianAvatar an efficient approach to creating realistic human avatars with dynamic 3D appearances from a single video. We start by introducing animatable 3D Gaussians to explicitly represent humans in various poses and clothing styles. Such an explicit and animatable representation can…

2024

HHMR: Holistic Hand Mesh Recovery by Enhancing the Multimodal Controllability of Graph Diffusion Models

CVPR 2024highlight

Recent years have witnessed a trend of the deep integration of the generation and reconstruction paradigms. In this paper we extend the ability of controllable generative models for a more comprehensive hand mesh recovery task: direct hand mesh generation inpainting reconstruction and fitting in a s…

Cited by 7SourcePDFScholar
2024

HumanNorm: Learning Normal Diffusion Model for High-quality and Realistic 3D Human Generation

CVPR 2024poster

Recent text-to-3D methods employing diffusion models have made significant advancements in 3D human generation. However these approaches face challenges due to the limitations of text-to-image diffusion models which lack an understanding of 3D structures. Consequently these methods struggle to achie…

2024

Learning Explicit Contact for Implicit Reconstruction of Hand-Held Objects from Monocular Images

AAAI 2024technical

Reconstructing hand-held objects from monocular RGB images is an appealing yet challenging task. In this task, contacts between hands and objects provide important cues for recovering the 3D geometry of the hand-held objects. Though recent works have employed implicit functions to achieve impressive…

2024

Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives

CVPR 2024poster

We propose Lodge a network capable of generating extremely long dance sequences conditioned on given music. We design Lodge as a two-stage coarse to fine diffusion architecture and propose the characteristic dance primitives that possess significant expressiveness as intermediate representations bet…

2024

ProxyCap: Real-time Monocular Full-body Capture in World Space via Human-Centric Proxy-to-Motion Learning

CVPR 2024poster

Learning-based approaches to monocular motion capture have recently shown promising results by learning to regress in a data-driven manner. However due to the challenges in data collection and network designs it remains challenging to achieve real-time full-body capture while being accurate in world…

Cited by 13SourcePDFScholar
2023

A Reachability-Based Spatio-Temporal Sampling Strategy for Kinodynamic Motion Planning

RA-L 2023

By limiting the planning domain to “ <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">L</i> <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sub> Informed Set”, some sampling-based motion planners (SBMP

Cited by 2SourceScholar
2023

CaPhy: Capturing Physical Properties for Animatable Human Avatars

ICCV 2023poster

We present CaPhy, a novel method for reconstructing animatable human avatars with realistic dynamic properties for clothing. Specifically, we aim for capturing the geometric and physical properties of the clothing from real observations. This allows us to apply novel poses to the human avatar with p…

Cited by 15PDFScholar
2023

CloSET: Modeling Clothed Humans on Continuous Surface With Explicit Template Decomposition

CVPR 2023poster

Creating animatable avatars from static scans requires the modeling of clothing deformations in different poses. Existing learning-based methods typically add pose-dependent deformations upon a minimally-clothed mesh template or a learned implicit template, which have limitations in capturing detail…

Cited by 28SourcePDFScholar
2023

Delving Deep into Pixel Alignment Feature for Accurate Multi-View Human Mesh Recovery

AAAI 2023technical

Regression-based methods have shown high efficiency and effectiveness for multi-view human mesh recovery. The key components of a typical regressor lie in the feature extraction of input views and the fusion of multi-view features. In this paper, we present Pixel-aligned Feedback Fusion (PaFF) for a…

2023

Leveraging Intrinsic Properties for Non-Rigid Garment Alignment

ICCV 2023poster

We address the problem of aligning real-world 3D data of garments, which benefits many applications such as texture learning, physical parameter estimation, generative modeling of garments, etc. Existing extrinsic methods typically perform non-rigid iterative closest point and struggle to align deta…

Cited by 7PDFcodeScholar
2023

Narrator: Towards Natural Control of Human-Scene Interaction Generation via Relationship Reasoning

ICCV 2023poster

Naturally controllable human-scene interaction (HSI) generation has an important role in various fields, such as VR/AR content creation and human-centered AI. However, existing methods are unnatural and unintuitive in their controllability, which heavily limits their application in practice. Therefo…

Cited by 9PDFScholar
2023

Next3D: Generative Neural Texture Rasterization for 3D-Aware Head Avatars

CVPR 2023highlight

3D-aware generative adversarial networks (GANs) synthesize high-fidelity and multi-view-consistent facial images using only collections of single-view 2D imagery. Towards fine-grained control over facial attributes, recent efforts incorporate 3D Morphable Face Model (3DMM) to describe deformation in…

2023

Tensor4D: Efficient Neural 4D Decomposition for High-Fidelity Dynamic Reconstruction and Rendering

CVPR 2023highlight

We present Tensor4D, an efficient yet effective approach to dynamic scene modeling. The key of our solution is an efficient 4D tensor decomposition method so that the dynamic scene can be directly represented as a 4D spatio-temporal tensor. To tackle the accompanying memory issue, we decompose the 4…

2022

AvatarCap: Animatable Avatar Conditioned Monocular Human Volumetric Capture

ECCV 2022poster

"To address the ill-posed problem caused by partial observations in monocular human volumetric capture, we present AvatarCap, a novel framework that introduces animatable avatars into the capture pipeline for high-fidelity reconstruction in both visible and invisible regions. Our method firstly crea…

2022

DiffuStereo: High Quality Human Reconstruction via Diffusion-Based Stereo Using Sparse Cameras

ECCV 2022poster

"We propose DiffuStereo, a novel system using only sparse cameras (8 in this work) for high-quality 3D human reconstruction. At its core is a novel diffusion-based stereo module, which introduces diffusion models, a type of powerful generative models, into the iterative stereo matching network. To t…

Cited by 68SourcePDFScholar
2022

DoubleField: Bridging the Neural Surface and Radiance Fields for High-Fidelity Human Reconstruction and Rendering

CVPR 2022poster

We introduce DoubleField, a novel framework combining the merits of both surface field and radiance field for high-fidelity human reconstruction and rendering. Within DoubleField, the surface field and radiance field are associated together by a shared feature embedding and a surface-guided sampling…

Cited by 186PDFScholar
2022

Interacting Attention Graph for Single Image Two-Hand Reconstruction

CVPR 2022oral

Graph convolutional network (GCN) has achieved great success in single hand reconstruction task, while interacting two-hand reconstruction by GCN remains unexplored. In this paper, we present Interacting Attention Graph Hand (IntagHand), the first graph convolution based network that reconstructs tw…

Cited by 133PDFcodeScholar
2022

Learning Implicit Templates for Point-Based Clothed Human Modeling

ECCV 2022poster

"We present FITE, a First-Implicit-Then-Explicit framework for modeling human avatars in clothing. Our framework first learns implicit surface templates representing the coarse clothing topology, and then employs the templates to guide the generation of point sets which further capture pose-dependen…

2022

Structured Local Radiance Fields for Human Avatar Modeling

CVPR 2022poster

It is extremely challenging to create an animatable clothed human avatar from RGB videos, especially for loose clothes due to the difficulties in motion modeling. To address this problem, we introduce a novel representation on the basis of recent neural scene rendering techniques. The core of our re…

Cited by 218PDFcodeScholar
2021

Evolving Search Space for Neural Architecture Search

ICCV 2021poster

Automation of neural architecture design has been a coveted alternative to human experts. Various search methods have been proposed aiming to find the optimal architecture in the search space. One would expect the search results to improve when the search space grows larger since it would potentiall…

Cited by 55PDFcodeScholar
2021

PyMAF: 3D Human Pose and Shape Regression With Pyramidal Mesh Alignment Feedback Loop

ICCV 2021poster

Regression-based methods have recently shown promising results in reconstructing human meshes from monocular images. By directly mapping raw pixels to model parameters, these methods can produce parametric models in a feed-forward manner via neural networks. However, minor deviation in parameters ma…

Cited by 386PDFcodeScholar
2020

Cheaper Pre-training Lunch: An Efficient Paradigm for Object Detection

ECCV 2020poster

In this paper, we propose a general and efficient pre-training paradigm, Montage pre-training, for object detection. Montage pre-training needs only the target detection dataset while taking only 1/4 computational resources compared to the widely adopted ImageNet pre-training. To build such an effic…

Cited by 23SourcePDFScholar
2020

Disentangling and Unifying Graph Convolutions for Skeleton-Based Action Recognition

CVPR 2020oral

Spatial-temporal graphs have been widely used by skeleton-based action recognition algorithms to model human action dynamics. To capture robust movement patterns from these graphs, long-range and multi-scale context aggregation and spatial-temporal dependency modeling are critical aspects of a power…

Cited by 1312PDFcodeScholar
2020

Rethinking Pseudo-LiDAR Representation

ECCV 2020poster

The recently proposed pseudo-LiDAR based 3D detectors greatly improves the benchmark of monocular/stereo 3D detection task. However, the underlying mechanism is still obscure to the research community. In this paper, we perform an in-depth investigation and observe that the pseudo-LiDAR representati…

2018

Learning a High Fidelity Pose Invariant Model for High-resolution Face Frontalization

NeurIPS 2018poster

Face frontalization refers to the process of synthesizing the frontal view of a face from a given profile. Due to self-occlusion and appearance distortion in the wild, it is extremely challenging to recover faithful results and preserve texture details in a high-resolution. This paper proposes a Hi…

Cited by 113SourcePDFScholar