ICASSP 2025accepted0 citations

Attention Disentanglement for Semantic Diffusion Modeling in Text-to-Image Generation

Hsiang-Chun Yu, Jen-Tzung Chien

Abstract

Text-to-image model has been recently improved to generate the semantically rich high-quality images by strengthening natural language processing via transformer in a stable diffusion process. However, the challenges are still remained in accurately rendering the objects, colors and compositions, and in precisely resolving the ambiguities in textual descriptions. This paper presents an attention disentanglment for semantic diffusion where the semantic consistency is enhanced for text-to-image generation. By utilizing the cross-attention maps in stable diffusion, this method is feasible to control the semantic features during generation process without the need for retraining. Additionally, a vision-language model is merged to implement this process to ensure that the generated images are closely aligned with the input prompts. This study further presents the attention disentanglement for semantic diffusion where the schemes of syntax parsing and information-theoretic learning are implemented to separate the semantic attributes and relations, thereby improving the accuracy and coherence of the generated images in the experiments.

BibTeX
@inproceedings{icassp2025_attentiondisenta,
  title = {Attention Disentanglement for Semantic Diffusion Modeling in Text-to-Image Generation},
  author = {Hsiang-Chun Yu and Jen-Tzung Chien},
  booktitle = {ICASSP 2025},
  year = {2025}
}