2025
Curriculum Direct Preference Optimization for Diffusion and Consistency Models
CVPR 2025poster
Direct Preference Optimization (DPO) has been proposed as an effective and efficient alternative to reinforcement learning from human feedback (RLHF). In this paper, we propose a novel and enhanced version of DPO based on curriculum learning for text-to-image generation. Our method is divided into t…