2026
HiPER: Hierarchical Plan–Execute RL for Multi-Turn LLM Agents
ICML 2026poster
Training LLMs as interactive agents for multi-turn decision-making remains challenging, particularly in long-horizon tasks with sparse and delayed rewards, where agents must execute extended sequences of actions before receiving meaningful feedback. Most existing reinforcement learning (RL) methods …