2026
Milestone-Guided Policy Learning for Long-Horizon Language Agents
ICML 2026poster
While long-horizon agentic tasks require language agents to perform dozens of sequential decisions, training such agents with reinforcement learning remains challenging. We identify two root causes: credit misattribution, where correct early actions are penalized due to terminal failures, and sample…