← Search

Haoxian Zhang

10 accepted papers

2026

3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation

CVPR 2026

Existing methods for human motion control in video generation typically rely on either 2D poses or explicit 3D parametric models (e.g., SMPL) as control signals. However, 2D poses rigidly bind motion to the driving viewpoint, precluding novel-view synthesis. Explicit 3D models, though structurally i

Cited by 0SourcecodeScholar
2026

From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping

ICML 2026poster

Audio-driven visual dubbing aims to synchronize a video's lip movements with new speech but is fundamentally challenged by the lack of ideal training data: paired videos differing only in lip motion. Existing methods circumvent this via mask-based inpainting. However, masking inevitably destroys spa…

Cited by 0SourceScholar
2025

Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained Control

ICLR 2025poster

Speech-driven 3D talking face method should offer both accurate lip synchronization and controllable expressions. Previous methods solely adopt discrete emotion labels to globally control expressions throughout sequences while limiting flexible fine-grained facial control within the spatiotemporal d…

Cited by 0SourcePDFScholar
2025

GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation

ICCV 2025poster

Creating high-quality, generalizable speech-driven 3D talking heads remains a persistent challenge. Previous methods achieve satisfactory results for fixed viewpoints and small-scale audio variations, but they struggle with large head rotations and out-of-distribution (OOD) audio. Moreover, they are…

Cited by 0SourcePDFScholar
2025

OmniSync: Towards Universal Lip Synchronization via Diffusion Transformers

NeurIPS 2025spotlight

Lip synchronization is the task of aligning a speaker’s lip movements in video with corresponding speech audio, and it is essential for creating realistic, expressive video content. However, existing methods often rely on reference frames and masked-frame inpainting, which limit their robustness to…

Cited by 0SourceScholar
2023

FFHQ-UV: Normalized Facial UV-Texture Dataset for 3D Face Reconstruction

CVPR 2023poster

We present a large-scale facial UV-texture dataset that contains over 50,000 high-quality texture UV-maps with even illuminations, neutral expressions, and cleaned facial regions, which are desired characteristics for rendering realistic 3D face models under different lighting conditions. The datase…

2022

"HVC-Net: Unifying Homography, Visibility, and Confidence Learning for Planar Object Tracking"

ECCV 2022poster

"Robust and accurate planar tracking over a whole video sequence is vitally important for many vision applications. The key to planar object tracking is to find object correspondences, modeled by homography, between the reference image and the tracked image. Existing methods tend to obtain wrong cor…

Cited by 9SourcePDFScholar
2022

REALY: Rethinking the Evaluation of 3D Face Reconstruction

ECCV 2022poster

"The evaluation of 3D face reconstruction results typically relies on a rigid shape alignment between the estimated 3D model and the ground-truth scan. We observe that aligning two shapes with different reference points can largely affect the evaluation results. This poses difficulties for precisely…

2021

Smoothing the Disentangled Latent Style Space for Unsupervised Image-to-Image Translation

CVPR 2021poster

Image-to-Image (I2I) multi-domain translation models are usually evaluated also using the quality of their semantic interpolation results. However, state-of-the-art models frequently show abrupt changes in the image appearance during interpolation, and usually perform poorly in interpolations across…

Cited by 58PDFScholar
2020

A Flexible Recurrent Residual Pyramid Network for Video Frame Interpolation

ECCV 2020poster

Video frame interpolation (VFI) aims at synthesizing new video frames in-between existing frames to generate smoother high frame rate videos. Current methods usually use the fixed pre-trained networks to generate interpolated-frames for different resolutions and scenes. However, the fixed pre-traine…

Cited by 46SourcePDFScholar