← Search

Zhentao Yu

4 accepted papers

2026

Harmony: Harmonizing Audio and Video Generation through Cross-Task Synergy

CVPR 2026

The synthesis of synchronized audio-visual content is a key challenge in generative AI, with open-source models facing challenges in robust audio-video alignment. Our analysis reveals that this issue is rooted in three fundamental challenges of the joint diffusion process: (1) Correspondence Drift,

Cited by 0SourcecodeScholar
2025

HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation

CVPR 2025poster

We introduce HunyuanPortrait, a diffusion-based condition control method that employs implicit representations for highly controllable and lifelike portrait animation. Given a single portrait image as an appearance reference and video clips as driving templates, HunyuanPortrait can animate the chara…

2025

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

NeurIPS 2025poster

Despite recent advances in video generation, existing models still lack fine-grained controllability, especially for multi-subject customization with consistent identity and interaction. In this paper, we propose PolyVivid, a multi-subject video customization framework that enables flexible and iden…

Cited by 0SourceScholar
2023

Talking Head Generation with Probabilistic Audio-to-Visual Diffusion Priors

ICCV 2023poster

We introduce a novel framework for one-shot audio-driven talking head generation. Unlike prior works that require additional driving sources for controlled synthesis in a deterministic manner, we instead sample all holistic lip-irrelevant facial motions (i.e. pose, expression, blink, gaze, etc.) to…

Cited by 41PDFScholar