2025
Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges
ACL 2025finding
Existing benchmarks that assess Language Models (LMs) as Language Agents (LAs) for tool use primarily focus on stateless, single-turn interactions or partial evaluations, such as tool selection in a single turn, overlooking the inherent stateful nature of interactions in multi-turn applications. To…