2026
Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO
ICML 2026poster
Group Relative Policy Optimization (GRPO) trains Chain-of-Thought reasoning with verifiable rewards, but estimating thought-level advantages without value functions often suffers from high variance. Although tree-style branching is used in practice to reduce the variance, it lacks a theoretical expl…