2026
InfoPO: Information-Driven Policy Optimization for User-Centric Agents
ICML 2026poster
Real-world user requests to LLM agents are often underspecified. Agents must interact to acquire missing information and make correct downstream decisions. However, current multi-turn GRPO-based methods often rely on trajectory-level reward computation, which leads to credit assignment problems and …