← Search

Fakhraddin Alwajih

7 accepted papers

2025

JAWAHER: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarking

NAACL 2025long

Recent advancements in instruction fine-tuning, alignment methods such as reinforcement learning from human feedback (RLHF), and optimization techniques like direct preference optimization (DPO), have significantly enhanced the adaptability of large language models (LLMs) to user preferences. Howeve…

2025

Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs

ACL 2025long

As large language models (LLMs) become increasingly integrated into daily life, ensuring their cultural sensitivity and inclusivity is paramount. We introduce PALM, a year-long community-driven project covering all 22 Arab countries. The dataset contains instruction–response pairs in both Modern Sta…

2025

Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset

EMNLP 2025

Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets. To address this gap, we introduce PEARL, a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding. Constructed through

2025

Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks

NAACL 2025findings

In this paper, we introduce Swan, a family of embedding models centred around the Arabic language, addressing both small-scale and large-scale use cases. Swan includes two variants: Swan-Small, based on ARBERTv2, and Swan-Large, built on ArMistral, a pretrained Arabic large language model. To evalua…

2024

Casablanca: Data and Models for Multidialectal Arabic Speech Recognition

EMNLP 2024main

In spite of the recent progress in speech processing, the majority of world languages and dialects remain uncovered. This situation only furthers an already wide technological divide, thereby hindering technological and socioeconomic inclusion. This challenge is largely due to the absence of dataset…

2024

Gazelle: An Instruction Dataset for Arabic Writing Assistance

EMNLP 2024finding

Writing has long been considered a hallmark of human intelligence and remains a pinnacle task for artificial intelligence (AI) due to the intricate cognitive processes involved. Recently, rapid advancements in generative AI, particularly through the development of Large Language Models (LLMs), have…

Cited by 0SourcePDFScholar
2024

Peacock: A Family of Arabic Multimodal Large Language Models and Benchmarks

ACL 2024long

Multimodal large language models (MLLMs) have proven effective in a wide range of tasks that require complex reasoning and linguistic comprehension. However, due to a lack of high-quality multimodal resources in languages other than English, the success of MLLMs remains relatively limited to English…