AAAI 2025technical4 citations

Generalizing Alignment Paradigm of Text-to-Image Generation with Preferences Through f-Divergence Minimization

Haoyuan Sun, Bo Xia, Yongzhe Chang, Xueqian Wang

Abstract

Direct Preference Optimization (DPO) has recently expanded its successful application from aligning large language models (LLMs) to aligning text-to-image models with human preferences, which has generated considerable interest within the community. However, we have observed that these approaches rely solely on minimizing the reverse Kullback-Leibler divergence during alignment process between the fine-tuned model and the reference model, neglecting incorporation of other divergence constraints. In this study, we focus on extending reverse Kullback-Leibler divergence in the alignment paradigm of text-to-image models to f-divergence, which aims to garner better alignment performance as well as good generation diversity. We provide the generalized formula of text-to-image alignment paradigm under f-divergence condition and thoroughly analyze the impact of different divergence constraints on alignment process from the perspective of gradient fields. We conduct comprehensive evaluation on text-image alignment performance, human value alignment performance and generation diversity performance under different divergence constraints, and the results indicate that text-to-image alignment based on Jensen-Shannon divergence achieves the best trade-off among them. The option of divergence employed for aligning text-to-image models significantly impacts the trade-off between alignment performance (especially human value alignment) and generation diversity, which highlights the necessity of selecting an appropriate divergence for practical applications.

BibTeX
@article{Sun_Xia_Chang_Wang_2025, title={Generalizing Alignment Paradigm of Text-to-Image Generation with Preferences Through f-Divergence Minimization}, volume={39}, url={https://ojs.aaai.org/index.php/AAAI/article/view/34978}, DOI={10.1609/aaai.v39i26.34978}, abstractNote={Direct Preference Optimization (DPO) has recently expanded its successful application from aligning large language models (LLMs) to aligning text-to-image models with human preferences, which has generated considerable interest within the community. However, we have observed that these approaches rely solely on minimizing the reverse Kullback-Leibler divergence during alignment process between the fine-tuned model and the reference model, neglecting incorporation of other divergence constraints. In this study, we focus on extending reverse Kullback-Leibler divergence in the alignment paradigm of text-to-image models to f-divergence, which aims to garner better alignment performance as well as good generation diversity. We provide the generalized formula of text-to-image alignment paradigm under f-divergence condition and thoroughly analyze the impact of different divergence constraints on alignment process from the perspective of gradient fields. We conduct comprehensive evaluation on text-image alignment performance, human value alignment performance and generation diversity performance under different divergence constraints, and the results indicate that text-to-image alignment based on Jensen-Shannon divergence achieves the best trade-off among them. The option of divergence employed for aligning text-to-image models significantly impacts the trade-off between alignment performance (especially human value alignment) and generation diversity, which highlights the necessity of selecting an appropriate divergence for practical applications.}, number={26}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Sun, Haoyuan and Xia, Bo and Chang, Yongzhe and Wang, Xueqian}, year={2025}, month={Apr.}, pages={27644-27652} }
Generalizing Alignment Paradigm of Text-to-Image Generation with Preferences Through f-Divergence Minimization · AAAI 2025