Semantic Attention and LLM-based Layout Guidance for Text-to-Image Generation
Yuxiang Song, Zhaoguang Long, Man Lan, Changzhi Sun, Aimin Zhou, Yuefeng Chen, Hao Yuan, Fei Cao
Abstract
Diffusion models have substantially advanced text-to-image generation, achieving remarkable performance in creating high-quality images from textual prompts. However, they often struggle with accurately generating images representing spatial locations described or implied in the prompts. To address this, we introduce SALT, a training-free method leveraging semantic attention and layout guidance from Large Language Models (LLMs) for text-to-image generation. This method effectively guides both cross-attention and self-attention layers within diffusion models, steering generation toward the direction of high-attention values provided by the layout guidance. During the denoising process of the diffusion model, image features in the latent space are iteratively refined based on the loss function calculated from the desired attention maps. Our approach has been executed on two benchmarks, providing detailed qualitative examples and comprehensive quantitative analyses. Results demonstrate that SALT outperforms existing training-free methods in controlling object layouts and generating attributes.<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>
BibTeX
@inproceedings{icassp2025_semanticattentio,
title = {Semantic Attention and LLM-based Layout Guidance for Text-to-Image Generation},
author = {Yuxiang Song and Zhaoguang Long and Man Lan and Changzhi Sun and Aimin Zhou and Yuefeng Chen and Hao Yuan and Fei Cao},
booktitle = {ICASSP 2025},
year = {2025}
}