2026
RiskPO: Risk-based Policy Optimization with Verifiable Reward for LLM Post-Training
ICLR 2026poster
Reinforcement learning with verifiable reward has recently emerged as a central paradigm for post-training large language models (LLMs); however, prevailing mean-based methods, such as Group Relative Policy Optimization (GRPO), suffer from entropy collapse and limited reasoning gains. We argue that…