2024
MaxFusion: Plug&Play Multi-Modal Generation in Text-to-Image Diffusion Models
ECCV 2024poster
"Large diffusion-based Text-to-Image (T2I) models have shown impressive generative powers for text-to-image generation and spatially conditioned image generation. We can train the model end-to-end with paired data for most applications to obtain photorealistic generation quality. However, to add a t…