2025
Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models
NeurIPS 2025poster
Enhancing the reasoning capabilities of large language models effectively using reinforcement learning (RL) remains a crucial challenge. Existing approaches primarily adopt two contrasting advantage estimation granularities: token-level methods (e.g., PPO) aim to provide fine-grained advantage signa…