← Search

Jinpeng Ou

2 accepted papers

2026

From Absolute to Relative: Rethinking Reward Shaping in Group-Based Reinforcement Learning

ICML 2026poster

Reinforcement learning has become a cornerstone for enhancing the reasoning capabilities of Large Language Models, where group-based approaches such as GRPO have emerged as efficient paradigms that optimize policies by leveraging intra-group performance differences. However, these methods typically …

Cited by 0SourceScholar
2026

THINK-AUGMENTED FUNCTION CALLING: IMPROVING LLM PARAMETER ACCURACY THROUGH EMBEDDED REASONING

ICASSP 2026poster

Large language models (LLMs) have demonstrated remarkable capabilities in function calling for autonomous agents, yet current mechanisms lack explicit reasoning transparency during parameter generation, particularly for complex functions with interdependent parameters. While existing approaches like…

Cited by 0SourcePDFScholar