← Search

Tamer Alkhouli

2 accepted papers

2025

CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions

ACL 2025long

We introduce Conversational Function-Calling Evaluation Through Turn-Level Interactions (CONFETTI), a conversational benchmark designed to evaluate the function-calling capabilities and response quality of large language models (LLMs). Current benchmarks lack comprehensive assessment of LLMs in comp…

2024

Eliciting Better Multilingual Structured Reasoning from LLMs through Code

ACL 2024long

The development of large language models (LLM) has shown progress on reasoning, though studies have largely considered either English or simple reasoning tasks. To address this, we introduce a multilingual structured reasoning and explanation dataset, termed xSTREET, that covers four tasks across si…

Cited by 8SourcePDFScholar