← Search

Zhihang Zheng

2 accepted papers

2026

GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy

ICML 2026poster

Reinforcement Learning (RL) is pivotal for enhancing Large Language Model (LLM) reasoning, yet mainstream algorithms such as GRPO and DAPO remain constrained by a coarse-grained credit assignment paradigm, where all tokens within the same response receive the identical reward. In this paper, we prop…

Cited by 0SourceScholar