2025
HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Assistant Scenarios
ACL 2025finding
Evaluating the performance of LLMs in multi-turn human-agent interactions presents significant challenges, particularly due to the complexity and variability of user behavior. In this paper, we introduce HammerBench, a novel benchmark framework for assessing LLMs’ function-calling capabilities in re…