← Search

Eman Albilali

1 accepted papers

2025

AraEval: An Arabic Multi-Task Evaluation Suite for Large Language Models

EMNLP 2025

The rapid advancements of Large Language models (LLMs) necessitate robust benchmarks. In this paper, we present AraEval, a pioneering and comprehensive evaluation suite specifically developed to assess the advanced knowledge, reasoning, truthfulness, and instruction- following capabilities of founda