2025
Entropy-Adaptive Diffusion Policy Optimization with Dynamic Step Alignment
ICCV 2025poster
While fine-tuning diffusion models with reinforcement learning (RL) has demonstrated effectiveness in directly optimizing downstream objectives, existing RL frameworks are prone to overfitting the rewards, leading to outputs that deviate from the true data distribution and exhibit reduced diversity.…