← Search

Jingyuan Yan

1 accepted papers

2026

GLARE: Scalable Neuro-Symbolic Reward Shaping for LLM Agents via Group-Level Automata

ICML 2026poster

Reinforcement Learning (RL) with Group Relative Policy Optimization (GRPO) shows great promise for enhancing LLM reasoning, but remains challenged by sparse and unstable rewards in long-horizon tasks. Existing approaches to reward shaping struggle to balance semantic expressiveness, reliability, and…

Cited by 0SourceScholar