← Search

Shamsuddeen Hassan Muhammad

13 accepted papers

2025

AFRIDOC-MT: Document-level MT Corpus for African Languages

EMNLP 2025

This paper introduces AFRIDOC-MT, a document-level multi-parallel translation dataset covering English and five African languages: Amharic, Hausa, Swahili, Yorùbá, and Zulu. The dataset comprises 334 health and 271 information technology news documents, all human-translated from English to these lan

2025

AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages

NAACL 2025long

Hate speech and abusive language are global phenomena that need socio-cultural background knowledge to be understood, identified, and moderated. However, in many regions of the Global South, there have been several documented occurrences of (1) absence of moderation and (2) censorship due to the rel…

2025

AfroXLMR-Social: Adapting Pre-trained Language Models for African Languages Social Media Text

EMNLP 2025

Language models built from various sources are the foundation of today’s NLP progress. However, for many low-resource languages, the diversity of domains is often limited, more biased to a religious domain, which impacts their performance when evaluated on distant and rapidly evolving domains such a

Cited by 0SourcePDFScholar
2025

BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages

ACL 2025long

People worldwide use language in subtle and complex ways to express emotions. Although emotion recognition–an umbrella term for several NLP tasks–impacts various applications within NLP and beyond, most work in this area has focused on high-resource languages. This has led to significant disparities…

2025

INJONGO: A Multicultural Intent Detection and Slot-filling Dataset for 16 African Languages

ACL 2025long

Slot-filling and intent detection are well-established tasks in Conversational AI. However, current large-scale benchmarks for these tasks often exclude evaluations of low-resource languages and rely on translations from English benchmarks, thereby predominantly reflecting Western-centric concepts.…

Cited by 0SourcePDFScholar
2025

IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models

NAACL 2025long

Despite the widespread adoption of Large language models (LLMs), their remarkable capabilities remain limited to a few high-resource languages. Additionally, many low-resource languages (e.g. African languages) are often evaluated only on basic text classification tasks due to the lack of appropriat…

2024

AfriMTE and AfriCOMET: Enhancing COMET to Embrace Under-resourced African Languages

NAACL 2024long

Despite the recent progress on scaling multilingual machine translation (MT) to several under-resourced African languages, accurately measuring this progress remains challenging, since evaluation is often performed on n-gram matching metrics such as BLEU, which typically show a weaker correlation wi…

2024

BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages

NeurIPS 2024poster

Large language models (LLMs) often lack culture-specific everyday knowledge, especially across diverse regions and non-English languages. Existing benchmarks for evaluating LLMs' cultural sensitivities are usually limited to a single language or online sources like Wikipedia, which may not reflect t…

2024

Mitigating Translationese in Low-resource Languages: The Storyboard Approach

COLING 2024main

Low-resource languages often face challenges in acquiring high-quality language data due to the reliance on translation-based methods, which can introduce the translationese effect. This phenomenon results in translated sentences that lack fluency and naturalness in the target language. In this pape…

Cited by 2SourcePDFScholar
2023

AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages

EMNLP 2023long main

Africa is home to over 2,000 languages from over six language families and has the highest linguistic diversity among all continents. This includes 75 languages with at least one million speakers each. Yet, there is little NLP research conducted on African languages. Crucial in enabling such researc…

Cited by 0SourcecodeScholar
2023

Cross-lingual Open-Retrieval Question Answering for African Languages

EMNLP 2023long findings

African languages have far less in-language content available digitally, making it challenging for question answering systems to satisfy the information needs of users. Cross-lingual open-retrieval question answering (XOR QA) systems -- those that retrieve answer content from other languages while s…

Cited by 0SourceScholar
2023

HaVQA: A Dataset for Visual Question Answering and Multimodal Research in Hausa Language

ACL 2023findings

This paper presents “HaVQA”, the first multimodal dataset for visual question answering (VQA) tasks in the Hausa language. The dataset was created by manually translating 6,022 English question-answer pairs, which are associated with 1,555 unique images from the Visual Genome dataset. As a result, t…

2023

MasakhaPOS: Part-of-Speech Tagging for Typologically Diverse African languages

ACL 2023long

In this paper, we present AfricaPOS, the largest part-of-speech (POS) dataset for 20 typologically diverse African languages. We discuss the challenges in annotating POS for these languages using the universal dependencies (UD) guidelines. We conducted extensive POS baseline experiments using both c…