Consistent Noisy Latent Rewards for Trajectory Preference Optimization in Diffusion Models
Recent advances in diffusion models for visual generation have sparked interest in human preference alignment, similar to developments in Large Language Models. While reward model (RM) based approaches enable trajectory-aware optimization by evaluating intermediate timesteps, they face two critical…