2024
Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model
CVPR 2024poster
Using reinforcement learning with human feedback (RLHF) has shown significant promise in fine-tuning diffusion models. Previous methods start by training a reward model that aligns with human preferences then leverage RL techniques to fine-tune the underlying models. However crafting an efficient re…