2026
Tool Call Dependency Graphs Enable Deep LLM Reasoning Evaluation and Better Explanations
IJCAI 2026
Evaluation of tool-augmented Large Language Models (LLMs) has not advanced far beyond final answer accuracy, and neglects in-depth evaluation of reasoning ability despite it being a central claim of recent models. We aim to address this gap by developing a dependency graph-based evaluation to give i