← Search

Yinuo Huang

1 accepted papers

2026

Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO

ICML 2026poster

Group Relative Policy Optimization (GRPO) trains Chain-of-Thought reasoning with verifiable rewards, but estimating thought-level advantages without value functions often suffers from high variance. Although tree-style branching is used in practice to reduce the variance, it lacks a theoretical expl…

Cited by 0SourceScholar