2025
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
NeurIPS 2025poster
Recent advancements underscore the significant role of Reinforcement Learning (RL) in enhancing the Chain-of-Thought (CoT) reasoning capabilities of large language models (LLMs). Two prominent RL algorithms, Direct Preference Optimization (DPO) and Group Relative Policy Optimization (GRPO), are cent…