← Search

Katherine Crowson

3 accepted papers

2024

Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion Transformers

ICML 2024poster

We present the Hourglass Diffusion Transformer (HDiT), an image-generative model that exhibits linear scaling with pixel count, supporting training at high resolution (e.g. $1024 \times 1024$) directly in pixel-space. Building on the Transformer architecture, which is known to scale to billions of p…

2022

LAION-5B: An open large-scale dataset for training next generation image-text models

NeurIPS 2022accept

Groundbreaking language-vision architectures like CLIP and DALL-E proved the utility of training on large amounts of noisy image-text data, without relying on expensive accurate labels used in standard vision unimodal supervised learning. The resulting models showed capabilities of strong text-guide…

2022

VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language Guidance

ECCV 2022poster

"Image generation and manipulation requires technical expertise to use, inhibiting adoption. Current methods rely heavily on training to a specific domain (e.g., only faces), manual work or algorithm tuning to latent vector discovery, and manual effort in mask selection to alter only a part of an im…