2026
Conditional Advantage Estimation for Reinforcement Learning in Large Reasoning Models
ICLR 2026poster
Reinforcement Learning with Verifiable Rewards (RLVR) for large language models (LLMs) has achieved remarkable progress in enhancing LLMs’ reasoning capabilities on tasks with clear correctness criteria, such as mathematical reasoning tasks. Several training metrics, such as entropy or response leng…