← Search

Safaa Taher Abdelfadil

3 accepted papers

2025

JAWAHER: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarking

NAACL 2025long

Recent advancements in instruction fine-tuning, alignment methods such as reinforcement learning from human feedback (RLHF), and optimization techniques like direct preference optimization (DPO), have significantly enhanced the adaptability of large language models (LLMs) to user preferences. Howeve…

2025

Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs

ACL 2025long

As large language models (LLMs) become increasingly integrated into daily life, ensuring their cultural sensitivity and inclusivity is paramount. We introduce PALM, a year-long community-driven project covering all 22 Arab countries. The dataset contains instruction–response pairs in both Modern Sta…

2025

Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset

EMNLP 2025

Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets. To address this gap, we introduce PEARL, a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding. Constructed through