2026
SSVPO: Effective Step-Level Credit Assignment for RL Training of Language Models
ICLR 2026poster
Language models have shown strong performance on mathematical reasoning tasks. Post-training with outcome-based reinforcement learning (RL) can further enhance reasoning but is inefficient because it relies solely on final rewards. Recent credit assignment–based RL methods provide intermediate feedb…