2026
Consistent Noisy Latent Rewards for Trajectory Preference Optimization in Diffusion Models
ICLR 2026poster
Recent advances in diffusion models for visual generation have sparked interest in human preference alignment, similar to developments in Large Language Models. While reward model (RM) based approaches enable trajectory-aware optimization by evaluating intermediate timesteps, they face two critical…