2026
Diffusion Blend: Inference-Time Multi-Preference Alignment for Diffusion Models
ICLR 2026poster
Reinforcement learning (RL) algorithms have been used recently to align diffusion models with downstream objectives such as aesthetic quality and text-image consistency by fine-tuning them to maximize a single reward function under a fixed KL regularization. However, this approach is inherently rest…