2025
Ranking-based Preference Optimization for Diffusion Models from Implicit User Feedback
NeurIPS 2025poster
Direct preference optimization (DPO) methods have shown strong potential in aligning text-to-image diffusion models with human preferences by training on paired comparisons. These methods improve training stability by avoiding the REINFORCE algorithm but still struggle with challenges such as accura…