← Search

Guozhen Zhang

13 accepted papers

2026

ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars

CVPR 2026

Despite significant advances in talking avatar generation, existing methods face critical challenges: insufficient text-following capability for diverse actions, lack of temporal alignment between actions and audio content, and dependency on additional control signals such as pose skeletons. We pres

Cited by 0SourceScholar
2026

Harmony: Harmonizing Audio and Video Generation through Cross-Task Synergy

CVPR 2026

The synthesis of synchronized audio-visual content is a key challenge in generative AI, with open-source models facing challenges in robust audio-video alignment. Our analysis reveals that this issue is rooted in three fundamental challenges of the joint diffusion process: (1) Correspondence Drift,

Cited by 0SourcecodeScholar
2026

StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human Avatars

CVPR 2026

Real-time, streaming interactive avatars represent a critical yet challenging goal in digital human research. Although diffusion-based human avatar generation methods achieve remarkable success, their non-causal architecture and high computational costs make them unsuitable for streaming. Moreover,

Cited by 0SourcecodeScholar
2026

UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions

CVPR 2026

Due to the lack of effective cross-modal modeling, existing open-source audio-video generation methods often exhibit compromised lip synchronization and insufficient semantic consistency. To mitigate these drawbacks, we propose UniAVGen, a unified framework for human-centric joint audio and video ge

Cited by 0SourceScholar
2025

Diffusion Transformers as Open-World Spatiotemporal Foundation Models

NeurIPS 2025poster

The urban environment is characterized by complex spatio-temporal dynamics arising from diverse human activities and interactions. Effectively modeling these dynamics is essential for understanding and optimizing urban systems. In this work, we introduce UrbanDiT, a foundation model for open-world u…

Cited by 0SourcecodeScholar
2025

OpenCarbon: A Contrastive Learning-based Cross-Modality Neural Approach for High-Resolution Carbon Emission Prediction Using Open Data

IJCAI 2025

Accurately estimating high-resolution carbon emissions is crucial for effective emission governance and mitigation planning. While conventional methods for precise carbon accounting are hindered by substantial data collection efforts, the rise of open data and advanced learning techniques offers a p

2024

Sparse Global Matching for Video Frame Interpolation with Large Motion

CVPR 2024poster

Large motion poses a critical challenge in Video Frame Interpolation (VFI) task. Existing methods are often constrained by limited receptive fields resulting in sub-optimal performance when handling scenarios with large motion. In this paper we introduce a new pipeline for VFI which can effectively…

Cited by 14SourcePDFScholar
2024

StableDrag: Stable Dragging for Point-based Image Editing

ECCV 2024poster

"Point-based image editing has attracted remarkable attention since the emergence of DragGAN. Recently, DragDiffusion further pushes forward the generative quality via adapting this dragging technique to diffusion models. Despite these great success, this dragging scheme exhibits two major drawbacks…

Cited by 12SourcePDFScholar
2024

VFIMamba: Video Frame Interpolation with State Space Models

NeurIPS 2024poster

Inter-frame modeling is pivotal in generating intermediate frames for video frame interpolation (VFI). Current approaches predominantly rely on convolution or attention-based models, which often either lack sufficient receptive fields or entail significant computational overheads. Recently, Selectiv…

2023

Extracting Motion and Appearance via Inter-Frame Attention for Efficient Video Frame Interpolation

CVPR 2023poster

Effectively extracting inter-frame motion and appearance information is important for video frame interpolation (VFI). Previous works either extract both types of information in a mixed way or devise separate modules for each type of information, which lead to representation ambiguity and low effici…

2023

MGMAE: Motion Guided Masking for Video Masked Autoencoding

ICCV 2023poster

Masked autoencoding has shown excellent performance on self-supervised video representation learning. Temporal redundancy has led to a high masking ratio and customized masking strategy in VideoMAE. In this paper, we aim to further improve the performance of video masked autoencoding by introducing…

Cited by 39PDFcodeScholar