← Search

Zhengguang Zhou

4 accepted papers

2026

ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars

CVPR 2026

Despite significant advances in talking avatar generation, existing methods face critical challenges: insufficient text-following capability for diverse actions, lack of temporal alignment between actions and audio content, and dependency on additional control signals such as pose skeletons. We pres

Cited by 0SourceScholar
2026

Harmony: Harmonizing Audio and Video Generation through Cross-Task Synergy

CVPR 2026

The synthesis of synchronized audio-visual content is a key challenge in generative AI, with open-source models facing challenges in robust audio-video alignment. Our analysis reveals that this issue is rooted in three fundamental challenges of the joint diffusion process: (1) Correspondence Drift,

Cited by 0SourcecodeScholar
2026

StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human Avatars

CVPR 2026

Real-time, streaming interactive avatars represent a critical yet challenging goal in digital human research. Although diffusion-based human avatar generation methods achieve remarkable success, their non-causal architecture and high computational costs make them unsuitable for streaming. Moreover,

Cited by 0SourcecodeScholar
2025

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

NeurIPS 2025poster

Despite recent advances in video generation, existing models still lack fine-grained controllability, especially for multi-subject customization with consistent identity and interaction. In this paper, we propose PolyVivid, a multi-subject video customization framework that enables flexible and iden…

Cited by 0SourceScholar