← Search

Andreas Blattmann

11 accepted papers

2024

SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

ICLR 2024spotlight

We present Stable Diffusion XL (SDXL), a latent diffusion model for text-to-image synthesis. Compared to previous versions of Stable Diffusion, SDXL leverages a three times larger UNet backbone, achieved by significantly increasing the number of attention blocks and including a second text encoder.…

2024

Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

ICML 2024oral

Diffusion models create data from noise by inverting the forward paths of data towards noise and have emerged as a powerful generative modeling technique for high-dimensional, perceptual data such as images and videos. Rectified flow is a recent generative model formulation that connects data and no…

Cited by 1056SourcePDFScholar
2023

Align Your Latents: High-Resolution Video Synthesis With Latent Diffusion Models

CVPR 2023poster

Latent Diffusion Models (LDMs) enable high-quality image synthesis while avoiding excessive compute demands by training a diffusion model in a compressed lower-dimensional latent space. Here, we apply the LDM paradigm to high-resolution video generation, a particularly resource-intensive task. We fi…

2022

High-Resolution Image Synthesis With Latent Diffusion Models

CVPR 2022oral

By decomposing the image formation process into a sequential application of denoising autoencoders, diffusion models (DMs) achieve state-of-the-art synthesis results on image data and beyond. Additionally, their formulation allows for a guiding mechanism to control the image generation process witho…

Cited by 18394PDFcodeScholar
2022

Retrieval-Augmented Diffusion Models

NeurIPS 2022accept

Novel architectures have recently improved generative image synthesis leading to excellent visual quality in various tasks. Much of this success is due to the scalability of these architectures and hence caused by a dramatic increase in model complexity and in the computational resources invested in…

Cited by 162SourcePDFScholar
2021

ImageBART: Bidirectional Context with Multinomial Diffusion for Autoregressive Image Synthesis

NeurIPS 2021poster

Autoregressive models and their sequential factorization of the data likelihood have recently demonstrated great potential for image representation and synthesis. Nevertheless, they incorporate image context in a linear 1D order by attending only to previously synthesized image patches above or to t…

Cited by 171SourcePDFScholar
2021

Stochastic Image-to-Video Synthesis Using cINNs

CVPR 2021poster

Video understanding calls for a model to learn the characteristic interplay between static scene content and its dynamics: Given an image, the model must be able to predict a future progression of the portrayed scene and, conversely, a video should be explained in terms of its static image content a…

Cited by 67PDFcodeScholar
2021

Understanding Object Dynamics for Interactive Image-to-Video Synthesis

CVPR 2021poster

What would be the effect of locally poking a static scene? We present an approach that learns naturally-looking global articulations caused by a local manipulation at a pixel level. Training requires only videos of moving objects but no information of the underlying manipulation of the physical scen…

Cited by 42PDFScholar
2021

iPOKE: Poking a Still Image for Controlled Stochastic Video Synthesis

ICCV 2021poster

How would a static scene react to a local poke? What are the effects on other parts of an object if you could locally push it? There will be distinctive movement, despite evident variations caused by the stochastic nature of our world. These outcomes are governed by the characteristic kinematics of…

Cited by 39PDFScholar