← Search

Ren Yangyang

1 accepted papers

2026

Unbiased Dynamic Pruning for Efficient Group-Based Policy Optimization

ICML 2026poster

Group Relative Policy Optimization (GRPO) effectively scales LLM reasoning but incurs prohibitive computational costs due to its extensive group-based sampling requirement. While recent selective data utilization methods can mitigate this overhead, they could induce estimation bias by altering the u…

Cited by 0SourceScholar