NeurIPS 2024poster7 citations

RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion Models

Xinchen Zhang, Ling Yang, YaQi Cai, Zhaochen Yu, Kai-Ni Wang, xie jiake, Ye Tian, Minkai Xu

Abstract

Diffusion models have achieved remarkable advancements in text-to-image generation. However, existing models still have many difficulties when faced with multiple-object compositional generation. In this paper, we propose ***RealCompo***, a new *training-free* and *transferred-friendly* text-to-image generation framework, which aims to leverage the respective advantages of text-to-image models and spatial-aware image diffusion models (e.g., layout, keypoints and segmentation maps) to enhance both realism and compositionality of the generated images. An intuitive and novel *balancer* is proposed to dynamically balance the strengths of the two models in denoising process, allowing plug-and-play use of any model without extra training. Extensive experiments show that our RealCompo consistently outperforms state-of-the-art text-to-image models and spatial-aware image diffusion models in multiple-object compositional generation while keeping satisfactory realism and compositionality of the generated images. Notably, our RealCompo can be seamlessly extended with a wide range of spatial-aware image diffusion models and stylized diffusion models. Code is available at: https://github.com/YangLing0818/RealCompo

Text-to-Image DiffusionLayout-guided Image Diffusion
BibTeX
@inproceedings{
zhang2024realcompo,
title={RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion Models},
author={Xinchen Zhang and Ling Yang and YaQi Cai and Zhaochen Yu and Kai-Ni Wang and xie jiake and Ye Tian and Minkai Xu and Yong Tang and Yujiu Yang and Bin CUI},
booktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},
year={2024},
url={https://openreview.net/forum?id=R8mfn3rHd5}
}
RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion Models · NeurIPS 2024