← Search

Abdulhamed Alothaimen

2 accepted papers

2025

AraEval: An Arabic Multi-Task Evaluation Suite for Large Language Models

EMNLP 2025

The rapid advancements of Large Language models (LLMs) necessitate robust benchmarks. In this paper, we present AraEval, a pioneering and comprehensive evaluation suite specifically developed to assess the advanced knowledge, reasoning, truthfulness, and instruction- following capabilities of founda

2025

LC-Eval: A Bilingual Multi-Task Evaluation Benchmark for Long-Context Understanding

EMNLP 2025

Recent advancements in Large Language Models (LLMs) have demonstrated sophisticated capabilities, including the ability to process and comprehend extended contexts. These emergent capabilities necessitate rigorous evaluation methods to effectively assess their performance in long-context understandi

Cited by 0SourcePDFScholar