ICASSP 2025accepted0 citations

Aligning Text-to-Image Diffusion Models without Human Feedback

Tao Liu, Huafeng Kuang, Xianming Lin

Abstract

Incorporating human feedback to optimize text-to-image models has demonstrated significant effectiveness. However, the process of collecting high-quality human preference labels is both resource-intensive and time-consuming. To address this challenge, we propose a novel approach that leverages a large language model (LLM) to generate sophisticated prompts, guiding the diffusion model towards enhanced image generation. This process inherently produces ranking pairs that approximate human preferences. We further introduce a novel integration of AI feedback with a Supervised Fine-Tuning (SFT) policy, aligning the model with preference labels derived from AI. Our experiments demonstrate that our approach achieves a notable approximation of human preferences, achieving a performance level of 68.13% compared to human-level benchmarks and delivering competitive results. Furthermore, we showcase the synergistic effects of combining AI feedback with human feedback, resulting in further improvements in image quality. This research offers fresh insights into AI feedback learning within text-to-image generation and lays the groundwork for more efficient and cost-effective training methodologies.

BibTeX
@inproceedings{icassp2025_aligningtexttoim,
  title = {Aligning Text-to-Image Diffusion Models without Human Feedback},
  author = {Tao Liu and Huafeng Kuang and Xianming Lin},
  booktitle = {ICASSP 2025},
  year = {2025}
}