← Search

Nora Al-Twairesh

4 accepted papers

2025

ALLaM: Large Language Models for Arabic and English

ICLR 2025poster

In this work, we present ALLaM: Arabic Large Language Model, a series of large language models to support the ecosystem of Arabic Language Technologies (ALT). ALLaM is carefully trained, considering the values of language alignment and transferability of knowledge at scale. The models are based on a…

Cited by 11SourcePDFScholar
2025

AraEval: An Arabic Multi-Task Evaluation Suite for Large Language Models

EMNLP 2025

The rapid advancements of Large Language models (LLMs) necessitate robust benchmarks. In this paper, we present AraEval, a pioneering and comprehensive evaluation suite specifically developed to assess the advanced knowledge, reasoning, truthfulness, and instruction- following capabilities of founda

2025

LC-Eval: A Bilingual Multi-Task Evaluation Benchmark for Long-Context Understanding

EMNLP 2025

Recent advancements in Large Language Models (LLMs) have demonstrated sophisticated capabilities, including the ability to process and comprehend extended contexts. These emergent capabilities necessitate rigorous evaluation methods to effectively assess their performance in long-context understandi

Cited by 0SourcePDFScholar
2024

When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

ACL 2024long

Large Language Model (LLM) leaderboards based on benchmark rankings are regularly used to guide practitioners in model selection. Often, the published leaderboard rankings are taken at face value — we show this is a (potentially costly) mistake. Under existing leaderboards, the relative performance…