ICASSP 2025accepted0 citations

Generative Adversarial Network with Structured Semantic Prompts Constrainting Clip for Text-to-Image

Shuheng Ge, Li Zhang, Haoyu Xing, Xiangqian Wu

Abstract

Rapidly synthesizing text-relevant images has long been a significant challenge. Introducing pre-trained models into GANs can significantly enhance model performance, enabling the fast generation of high-quality images. Previous work has primarily focused on improving visual quality, with limited fine-grained research on text-image consistency. This often leads to issues such as the loss of details and concept confusion in the synthesized images. To address these challenges, we propose a novel image synthesis method combining GANs with pre-trained models. This method captures fine-grained visual concepts from text and constructs structured semantic prompts, which hierarchically guide the adjustment of visual features in the latent space. It improves image quality while significantly enhancing text-image consistency. Moreover, we introduce hard mining into GAN-based image generation tasks and propose a hard mining matching-aware loss, enabling the model to focus on the most challenging samples. This reduces the model’s dependency on batch size and training epochs, achieving good performance even under limited computational resources. Extensive experiments validate the effectiveness of our approach.

BibTeX
@inproceedings{icassp2025_generativeadvers,
  title = {Generative Adversarial Network with Structured Semantic Prompts Constrainting Clip for Text-to-Image},
  author = {Shuheng Ge and Li Zhang and Haoyu Xing and Xiangqian Wu},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Generative Adversarial Network with Structured Semantic Prompts Constrainting Clip for Text-to-Image · ICASSP 2025