← Search

Zheng Ding

11 accepted papers

2024

Dolfin: Diffusion Layout Transformers without Autoencoder

ECCV 2024poster

"In this paper, we introduce a new generative model, Diffusion Layout Transformers without Autoencoder (Dolfin), that attains significantly improved modeling capability and transparency over the existing approaches. Dolfin employs a Transformer-based diffusion process to model layout generation. In…

Cited by 17SourcePDFScholar
2024

Explorative Inbetweening of Time and Space

ECCV 2024poster

"We introduce bounded generation as a generalized task to control video generation to synthesize arbitrary camera and subject motion based only on a given start and end frame. Our objective is to fully leverage the inherent generalization capability of an image-to-video model without additional trai…

Cited by 12SourcePDFScholar
2024

HOIDiffusion: Generating Realistic 3D Hand-Object Interaction Data

CVPR 2024poster

3D hand-object interaction data is scarce due to the hardware constraints in scaling up the data collection process. In this paper we propose HOIDiffusion for generating realistic and diverse 3D hand-object interaction data. Our model is a conditional diffusion model that takes both the 3D hand-obje…

2024

Patched Denoising Diffusion Models For High-Resolution Image Synthesis

ICLR 2024poster

We propose an effective denoising diffusion model for generating high-resolution images (e.g., 1024$\times$512), trained on small-size image patches (e.g., 64$\times$64). We name our algorithm Patch-DM, in which a new feature collage strategy is designed to avoid the boundary artifact when synthesiz…

2024

TokenCompose: Text-to-Image Diffusion with Token-level Supervision

CVPR 2024poster

We present TokenCompose a Latent Diffusion Model for text-to-image generation that achieves enhanced consistency between user-specified text prompts and model-generated images. Despite its tremendous success the standard denoising process in the Latent Diffusion Model takes text prompts as condition…

2023

DiffusionRig: Learning Personalized Priors for Facial Appearance Editing

CVPR 2023poster

We address the problem of learning person-specific facial priors from a small number (e.g., 20) of portrait photos of the same person. This enables us to edit this specific person's facial appearance, such as expression and lighting, while preserving their identity and high-frequency facial details.…

2020

Guided Variational Autoencoder for Disentanglement Learning

CVPR 2020poster

We propose an algorithm, guided variational autoencoder (Guided-VAE), that is able to learn a controllable generative model by performing latent representation disentanglement learning. The learning objective is achieved by providing signal to the latent encoding/embedding in VAE without changing it…

Cited by 151PDFcodeScholar