2026
Beyond Normalization: Rethinking the Partition Function as a Difficulty Scheduler for RLVR
ICML 2026poster
Reward-maximizing RL methods enhance the reasoning performance of LLMs, but often reduce the diversity among outputs. Recent works address this issue by adopting GFlowNets, training LLMs to match a target distribution while jointly learning its partition function. In contrast to prior works that tre…