2026
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
ICML 2026poster
Recent studies observe that reinforcement learning with verifiable rewards (RLVR) reliably improves pass@1 on reasoning tasks, yet often fails to yield comparable gains in pass@k, raising the question of whether RLVR genuinely enables large language models to acquire novel reasoning abilities or mer…