← Search

Omer Nacar

3 accepted papers

2025

NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communities

EMNLP 2025

Enhancing the linguistic capabilities of Large Language Models (LLMs) to include low-resource languages is a critical research area. Current research directions predominantly rely on synthetic data generated by translating English corpora, which, while demonstrating promising linguistic understandin

2025

Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs

ACL 2025long

As large language models (LLMs) become increasingly integrated into daily life, ensuring their cultural sensitivity and inclusivity is paramount. We introduce PALM, a year-long community-driven project covering all 22 Arab countries. The dataset contains instruction–response pairs in both Modern Sta…

2025

Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset

EMNLP 2025

Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets. To address this gap, we introduce PEARL, a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding. Constructed through