2026
PCPO: Proportionate Credit Policy Optimization for Preference Alignment of Image Generation Models
ICLR 2026poster
While reinforcement learning has advanced the alignment of text-to-image (T2I) models, state-of-the-art policy gradient methods are still hampered by training instability and high variance, hindering convergence speed and compromising image quality. Our analysis identifies a key cause of this instab…