← Search

Pengze Zhang

5 accepted papers

2026

DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation

ICML 2026poster

Recent advancements in foundation models have revolutionized joint audio-video generation. However, existing approaches typically treat human-centric tasks including reference-based audio-video generation (R2AV), video editing (RV2AV) and audio-driven video animation (RA2V) as isolated objectives. F…

Cited by 0SourceScholar
2025

AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models

CVPR 2025poster

Recent advances in garment-centric image generation from text and image prompts based on diffusion models are impressive. However, existing methods lack support for various combinations of attire, and struggle to preserve the garment details while maintaining faithfulness to the text prompts, limiti…

Cited by 5SourcePDFScholar
2024

Tackling the Singularities at the Endpoints of Time Intervals in Diffusion Models

CVPR 2024highlight

Most diffusion models assume that the reverse process adheres to a Gaussian distribution. However this approximation has not been rigorously validated especially at singularities where t=0 and t=1. Improperly dealing with such singularities leads to an average brightness issue in applications and li…

2023

Formulating Discrete Probability Flow Through Optimal Transport

NeurIPS 2023poster

Continuous diffusion models are commonly acknowledged to display a deterministic probability flow, whereas discrete diffusion models do not. In this paper, we aim to establish the fundamental theory for the probability flow of discrete diffusion models. Specifically, we first prove that the continuo…

2022

Exploring Dual-Task Correlation for Pose Guided Person Image Generation

CVPR 2022poster

Pose Guided Person Image Generation (PGPIG) is the task of transforming a person image from the source pose to a given target pose. Most of the existing methods only focus on the ill-posed source-to-target task and fail to capture reasonable texture mapping. To address this problem, we propose a nov…

Cited by 102PDFcodeScholar