2025
ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning
NeurIPS 2025poster
Recent advances in large language models have been driven by reinforcement learning (RL)-style post-training, which improves reasoning by optimizing model outputs based on reward or preference signals. GRPO-style approaches implement this by using self-generated samples labeled by an outcome-based v…