← Search

Hanfeng Zhao

4 accepted papers

2026

Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback

AAAI 2026technical

Recent advances in diffusion models have significantly improved audio-driven human video generation, surpassing traditional methods in both quality and controllability. However, existing approaches still face challenges in lip-sync accuracy, temporal coherence for long video generation, and multi-ch

Cited by 0SourcePDFScholar
2025

FlexGen: Flexible Multi-View Generation from Text and Image Inputs

ICCV 2025poster

In this work, we introduce FlexGen, a flexible framework designed to generate controllable and consistent multi-view images, conditioned on a single-view image, or a text prompt, or both. FlexGen tackles the challenges of controllable multi-view synthesis through additional conditioning on 3D-aware…

Cited by 0SourcePDFScholar
2025

GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs

ICCV 2025poster

Estimating physical properties for visual data is a crucial task in computer vision, graphics, and robotics, underpinning applications such as augmented reality, physical simulation, and robotic grasping. However, this area remains under-explored due to the inherent ambiguities in physical property…

Cited by 0SourcePDFScholar
2025

MultiGO: Towards Multi-level Geometry Learning for Monocular 3D Textured Human Reconstruction

CVPR 2025poster

This paper investigates the research task of reconstructing the 3D clothed human body from a monocular image. Due to the inherent ambiguity of single-view input, existing approaches leverage pre-trained SMPL(-X) estimation models or generative models to provide auxiliary information for human recons…

Cited by 3SourcePDFScholar