2026
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
ICML 2026poster
Reinforcement Learning (RL) is pivotal for enhancing Large Language Model (LLM) reasoning, yet mainstream algorithms such as GRPO and DAPO remain constrained by a coarse-grained credit assignment paradigm, where all tokens within the same response receive the identical reward. In this paper, we prop…