2025
CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions
ACL 2025long
We introduce Conversational Function-Calling Evaluation Through Turn-Level Interactions (CONFETTI), a conversational benchmark designed to evaluate the function-calling capabilities and response quality of large language models (LLMs). Current benchmarks lack comprehensive assessment of LLMs in comp…