← Search

Muhammad Abdul-Mageed

39 accepted papers

2025

Beyond Content: How Grammatical Gender Shapes Visual Representation in Text-to-Image Models

EMNLP 2025

Research on bias in Text-to-Image (T2I) models has primarily focused on demographic representation and stereotypical attributes, overlooking a fundamental question: how does grammatical gender influence visual representation across languages? We introduce a cross-linguistic benchmark examining words

Cited by 0SourcePDFScholar
2025

EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs

EMNLP 2025

Large language models (LLMs) are transforming education by answering questions, explaining complex concepts, and generating content across a wide range of subjects. Despite strong performance on academic benchmarks, they often fail to tailor responses to students’ grade levels. This is a critical ne

2025

Effective Self-Mining of In-Context Examples for Unsupervised Machine Translation with LLMs

NAACL 2025findings

Large Language Models (LLMs) have demonstrated impressive performance on a wide range of natural language processing (NLP) tasks, primarily through in-context learning (ICL). In ICL, the LLM is provided with examples that represent a given task such that it learns to generate answers for test inputs…

2025

JAWAHER: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarking

NAACL 2025long

Recent advancements in instruction fine-tuning, alignment methods such as reinforcement learning from human feedback (RLHF), and optimization techniques like direct preference optimization (DPO), have significantly enhanced the adaptability of large language models (LLMs) to user preferences. Howeve…

2025

NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communities

EMNLP 2025

Enhancing the linguistic capabilities of Large Language Models (LLMs) to include low-resource languages is a critical research area. Current research directions predominantly rely on synthetic data generated by translating English corpora, which, while demonstrating promising linguistic understandin

2025

Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs

ACL 2025long

As large language models (LLMs) become increasingly integrated into daily life, ensuring their cultural sensitivity and inclusivity is paramount. We introduce PALM, a year-long community-driven project covering all 22 Arab countries. The dataset contains instruction–response pairs in both Modern Sta…

2025

Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset

EMNLP 2025

Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets. To address this gap, we introduce PEARL, a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding. Constructed through

2025

Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks

NAACL 2025findings

In this paper, we introduce Swan, a family of embedding models centred around the Arabic language, addressing both small-scale and large-scale use cases. Swan includes two variants: Swan-Small, based on ARBERTv2, and Swan-Large, built on ArMistral, a pretrained Arabic large language model. To evalua…

2025

Voice of a Continent: Mapping Africa’s Speech Technology Frontier

EMNLP 2025

Africa’s rich linguistic diversity remains significantly underrepresented in speech technologies, creating barriers to digital inclusion. To alleviate this challenge, we systematically map the continent’s speech space of datasets and technologies, leading to a new comprehensive benchmark SimbaBench

2025

Where Are We? Evaluating LLM Performance on African Languages

ACL 2025long

Africa’s rich linguistic heritage remains underrepresented in NLP, largely due to historical policies that favor foreign languages and create significant data inequities. In this paper, we integrate theoretical insights on Africa’s language landscape with an empirical evaluation using Sahara— a comp…

Cited by 0SourcePDFScholar
2025

uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes

NAACL 2025long

Recent work on distilling Whisper’s knowledge into small models using pseudo-labels shows promising performance while reducing the size by up to 50%. This results in small, efficient, and dedicated models. However, a critical step of distillation using pseudo-labels involves filtering high-quality p…

2024

Casablanca: Data and Models for Multidialectal Arabic Speech Recognition

EMNLP 2024main

In spite of the recent progress in speech processing, the majority of world languages and dialects remain uncovered. This situation only furthers an already wide technological divide, thereby hindering technological and socioeconomic inclusion. This challenge is largely due to the absence of dataset…

2024

Cheetah: Natural Language Generation for 517 African Languages

ACL 2024long

Low-resource African languages pose unique challenges for natural language processing (NLP) tasks, including natural language generation (NLG). In this paper, we develop Cheetah, a massively multilingual NLG language model for African languages. Cheetah supports 517 African languages and language va…

2024

DetoxLLM: A Framework for Detoxification with Explanations

EMNLP 2024main

Prior works on detoxification are scattered in the sense that they do not cover all aspects of detoxification needed in a real-world scenario. Notably, prior works restrict the task of developing detoxification models to only a seen subset of platforms, leaving the question of how the models would p…

2024

FinTral: A Family of GPT-4 Level Multimodal Financial Large Language Models

ACL 2024findings

We introduce FinTral, a suite of state-of-the-art multimodal large language models (LLMs) built upon the Mistral-7b model and tailored for financial analysis. FinTral integrates textual, numerical, tabular, and image data. We enhance FinTral with domain-specific pretraining, instruction fine-tuning,…

2024

Fumbling in Babel: An Investigation into ChatGPT’s Language Identification Ability

NAACL 2024findings

ChatGPT has recently emerged as a powerful NLP tool that can carry out a variety of tasks. However, the range of languages ChatGPT can handle remains largely a mystery. To uncover which languages ChatGPT ‘knows’, we investigate its language identification (LID) abilities. For this purpose, we compil…

2024

Gazelle: An Instruction Dataset for Arabic Writing Assistance

EMNLP 2024finding

Writing has long been considered a hallmark of human intelligence and remains a pinnacle task for artificial intelligence (AI) due to the intricate cognitive processes involved. Recently, rapid advancements in generative AI, particularly through the development of Large Language Models (LLMs), have…

Cited by 0SourcePDFScholar
2024

LLM Performance Predictors are good initializers for Architecture Search

ACL 2024findings

In this work, we utilize Large Language Models (LLMs) for a novel use case: constructing Performance Predictors (PP) that estimate the performance of specific deep neural network architectures on downstream tasks. We create PP prompts for LLMs, comprising (i) role descriptions, (ii) instructions for…

2024

Mixture-of-Supernets: Improving Weight-Sharing Supernet Training with Architecture-Routed Mixture-of-Experts

ACL 2024findings

Weight-sharing supernets are crucial for performance estimation in cutting-edge neural architecture search (NAS) frameworks. Despite their ability to generate diverse subnetworks without retraining, the quality of these subnetworks is not guaranteed due to weight sharing. In NLP tasks like machine t…

2024

Peacock: A Family of Arabic Multimodal Large Language Models and Benchmarks

ACL 2024long

Multimodal large language models (MLLMs) have proven effective in a wide range of tasks that require complex reasoning and linguistic comprehension. However, due to a lack of high-quality multimodal resources in languages other than English, the success of MLLMs remains relatively limited to English…

2024

To Distill or Not to Distill? On the Robustness of Robust Knowledge Distillation

ACL 2024long

Arabic is known to present unique challengesfor Automatic Speech Recognition (ASR). Onone hand, its rich linguistic diversity andwide range of dialects complicate the de-velopment of robust, inclusive models. Onthe other, current multilingual ASR modelsare compute-intensive and lack proper com-prehe…

2024

Toucan: Many-to-Many Translation for 150 African Language Pairs

ACL 2024findings

We address a notable gap in Natural Language Processing (NLP) by introducing a collection of resources designed to improve Machine Translation (MT) for low-resource languages, with a specific focus on African languages. First, We introduce two language models (LMs), Cheetah-1.2B and Cheetah-3.7B, wi…

2023

AutoMoE: Heterogeneous Mixture-of-Experts with Adaptive Computation for Efficient Neural Machine Translation

ACL 2023findings

Mixture-of-Expert (MoE) models have obtained state-of-the-art performance in Neural Machine Translation (NMT) tasks. Existing works in MoE mostly consider a homogeneous design where the same number of experts of the same size are placed uniformly throughout the network. Furthermore, existing MoE wor…

2023

Contrastive Learning of Sociopragmatic Meaning in Social Media

ACL 2023findings

Recent progress in representation and contrastive learning in NLP has not widely considered the class of sociopragmatic meaning (i.e., meaning in interaction within different language communities). To bridge this gap, we propose a novel framework for learning task-agnostic representations transferab…

2023

Dolphin: A Challenging and Diverse Benchmark for Arabic NLG

EMNLP 2023long findings

We present Dolphin, a novel benchmark that addresses the need for a natural language generation (NLG) evaluation framework dedicated to the wide collection of Arabic languages and varieties. The proposed benchmark encompasses a broad range of 13 different NLG tasks, including dialogue generation, qu…

Cited by 0SourceScholar
2023

GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP

EMNLP 2023long main

ChatGPT's emergence heralds a transformative phase in NLP, particularly demonstrated through its excellent performance on many English benchmarks. However, the model's efficacy across diverse linguistic contexts remains largely uncharted territory. This work aims to bridge this knowledge gap, with a…

Cited by 0SourceScholar
2023

JASMINE: Arabic GPT Models for Few-Shot Learning

EMNLP 2023long main

Scholarship on generative pretraining (GPT) remains acutely Anglocentric, leaving serious gaps in our understanding of the whole class of autoregressive models. For example, we have little knowledge about the potential of these models and their societal impacts in diverse linguistic and cultural set…

Cited by 0SourceScholar
2023

ORCA: A Challenging Benchmark for Arabic Language Understanding

ACL 2023findings

Due to the crucial role pretrained language models play in modern NLP, several benchmarks have been proposed to evaluate their performance. In spite of these efforts, no public benchmark of diverse nature currently exists for evaluating Arabic NLU. This makes it challenging to measure progress for b…

Cited by 27SourcePDFScholar
2023

SERENGETI: Massively Multilingual Language Models for Africa

ACL 2023findings

Multilingual pretrained language models (mPLMs) acquire valuable, generalizable linguistic information during pretraining and have advanced the state of the art on task-specific finetuning. To date, only ~31 out of ~2,000 African languages are covered in existing language models. We ameliorate this…

2023

The Skipped Beat: A Study of Sociopragmatic Understanding in LLMs for 64 Languages

EMNLP 2023long main

Instruction tuned large language models (LLMs), such as ChatGPT, demonstrate remarkable performance in a wide range of tasks. Despite numerous recent studies that examine the performance of instruction-tuned LLMs on various NLP benchmarks, there remains a lack of comprehensive investigation into the…

Cited by 0SourcecodeScholar
2022

AfroLID: A Neural Language Identification Tool for African Languages

EMNLP 2022main

Language identification (LID) is a crucial precursor for NLP, especially for mining web data. Problematically, most of the world’s 7000+ languages today are not covered by LID technologies. We address this pressing issue for Africa by introducing AfroLID, a neural LID toolkit for 517 African languag…

2022

AraT5: Text-to-Text Transformers for Arabic Language Generation

ACL 2022long

Transfer learning with a unified Transformer framework (T5) that converts all language problems into a text-to-text format was recently proposed as a simple and effective transfer learning approach. Although a multilingual version of the T5 model (mT5) was also introduced, it is not clear how well i…

2022

Automatic Detection of Entity-Manipulated Text using Factual Knowledge

ACL 2022short

In this work, we focus on the problem of distinguishing a human written news article from a news article that is created by manipulating entities in a human written news article (e.g., replacing entities with factually incorrect entities). Such manipulated articles can mislead the reader by posing a…

2022

Linguistically-Motivated Yorùbá-English Machine Translation

COLING 2022main

Translating between languages where certain features are marked morphologically in one but absent or marked contextually in the other is an important test case for machine translation. When translating into English which marks (in)definiteness morphologically, from Yorùbá which uses bare nouns but m…

2022

Towards Afrocentric NLP for African Languages: Where We Are and Where We Can Go

ACL 2022long

Aligning with ACL 2022 special Theme on “Language Diversity: from Low Resource to Endangered Languages”, we discuss the major linguistic and sociopolitical challenges facing development of NLP technologies for African languages. Situating African languages in a typological framework, we discuss how…

Cited by 0SourcePDFScholar
2021

ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic

ACL 2021long

Pre-trained language models (LMs) are currently integral to many natural language processing systems. Although multilingual LMs were also introduced to serve many languages, these have limitations such as being costly at inference time and the size and diversity of non-English data involved in their…

2020

Automatic Detection of Machine Generated Text: A Critical Survey

COLING 2020main

Text generative models (TGMs) excel in producing text that matches the style of human language reasonably well. Such TGMs can be misused by adversaries, e.g., by automatically generating fake news and fake product reviews that can look authentic and fool humans. Detectors that can distinguish text g…

2019

Deep Learning the EEG Manifold for Phonological Categorization from Active Thoughts

ICASSP 2019accepted

Speech-related Brain Computer Interfaces (BCI) aim primarily at finding an alternative vocal communication pathway for people with speaking disabilities. As a step towards full decoding of imagined speech from active thoughts, we present a BCI system for subject-independent classification of phonolo…

Cited by 0SourceScholar