Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
Driven by recent advancements in tool-augmented Large Language Model (LLM) agents, comprehensive benchmark datasets for evaluating these tool-augmented agents are being actively developed. Although these benchmarks incorporate increasingly complex user requests and a diverse array of tools, the eval…