← Search

El Moatez Billah Nagoudi

10 accepted papers

2025

Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs

ACL 2025long

As large language models (LLMs) become increasingly integrated into daily life, ensuring their cultural sensitivity and inclusivity is paramount. We introduce PALM, a year-long community-driven project covering all 22 Arab countries. The dataset contains instruction–response pairs in both Modern Sta…

2025

Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks

NAACL 2025findings

In this paper, we introduce Swan, a family of embedding models centred around the Arabic language, addressing both small-scale and large-scale use cases. Swan includes two variants: Swan-Small, based on ARBERTv2, and Swan-Large, built on ArMistral, a pretrained Arabic large language model. To evalua…

2024

Casablanca: Data and Models for Multidialectal Arabic Speech Recognition

EMNLP 2024main

In spite of the recent progress in speech processing, the majority of world languages and dialects remain uncovered. This situation only furthers an already wide technological divide, thereby hindering technological and socioeconomic inclusion. This challenge is largely due to the absence of dataset…

2024

FinTral: A Family of GPT-4 Level Multimodal Financial Large Language Models

ACL 2024findings

We introduce FinTral, a suite of state-of-the-art multimodal large language models (LLMs) built upon the Mistral-7b model and tailored for financial analysis. FinTral integrates textual, numerical, tabular, and image data. We enhance FinTral with domain-specific pretraining, instruction fine-tuning,…

2024

Peacock: A Family of Arabic Multimodal Large Language Models and Benchmarks

ACL 2024long

Multimodal large language models (MLLMs) have proven effective in a wide range of tasks that require complex reasoning and linguistic comprehension. However, due to a lack of high-quality multimodal resources in languages other than English, the success of MLLMs remains relatively limited to English…

2023

Dolphin: A Challenging and Diverse Benchmark for Arabic NLG

EMNLP 2023long findings

We present Dolphin, a novel benchmark that addresses the need for a natural language generation (NLG) evaluation framework dedicated to the wide collection of Arabic languages and varieties. The proposed benchmark encompasses a broad range of 13 different NLG tasks, including dialogue generation, qu…

Cited by 0SourceScholar
2023

GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP

EMNLP 2023long main

ChatGPT's emergence heralds a transformative phase in NLP, particularly demonstrated through its excellent performance on many English benchmarks. However, the model's efficacy across diverse linguistic contexts remains largely uncharted territory. This work aims to bridge this knowledge gap, with a…

Cited by 0SourceScholar
2023

JASMINE: Arabic GPT Models for Few-Shot Learning

EMNLP 2023long main

Scholarship on generative pretraining (GPT) remains acutely Anglocentric, leaving serious gaps in our understanding of the whole class of autoregressive models. For example, we have little knowledge about the potential of these models and their societal impacts in diverse linguistic and cultural set…

Cited by 0SourceScholar
2022

AraT5: Text-to-Text Transformers for Arabic Language Generation

ACL 2022long

Transfer learning with a unified Transformer framework (T5) that converts all language problems into a text-to-text format was recently proposed as a simple and effective transfer learning approach. Although a multilingual version of the T5 model (mT5) was also introduced, it is not clear how well i…

2021

ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic

ACL 2021long

Pre-trained language models (LMs) are currently integral to many natural language processing systems. Although multilingual LMs were also introduced to serve many languages, these have limitations such as being costly at inference time and the size and diversity of non-English data involved in their…