2025
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
ICML 2025poster
Large Language Models excel at code generation yet struggle with complex programming tasks that demand sophisticated reasoning. To bridge this gap, traditional process supervision relies on learned reward models requiring costly training data and suffering from reward misalignment, while outcome su…