2026
Learn More with Less: Uncertainty Consistency Guided Query Selection for RLVR
ICLR 2026poster
Large Language Models (LLMs) have recently improved mathematical reasoning through Reinforcement Learning with Verifiable Reward (RLVR). However, existing RLVR algorithms require large query budgets, making annotation costly. We investigate whether fewer but more informative queries can yield simila…