← Search

Ibrahim Said Ahmad

9 accepted papers

2025

AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages

NAACL 2025long

Hate speech and abusive language are global phenomena that need socio-cultural background knowledge to be understood, identified, and moderated. However, in many regions of the Global South, there have been several documented occurrences of (1) absence of moderation and (2) censorship due to the rel…

2025

AfroXLMR-Social: Adapting Pre-trained Language Models for African Languages Social Media Text

EMNLP 2025

Language models built from various sources are the foundation of today’s NLP progress. However, for many low-resource languages, the diversity of domains is often limited, more biased to a religious domain, which impacts their performance when evaluated on distant and rapidly evolving domains such a

Cited by 0SourcePDFScholar
2025

BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages

ACL 2025long

People worldwide use language in subtle and complex ways to express emotions. Although emotion recognition–an umbrella term for several NLP tasks–impacts various applications within NLP and beyond, most work in this area has focused on high-resource languages. This has led to significant disparities…

2025

Is Peer-Reviewing Worth the Effort?

COLING 2025main

How effective is peer-reviewing in identifying important papers? We treat this question as a forecasting task. Can we predict which papers will be highly cited in the future based on venue and “early returns” (citations soon after publication)? We show early returns are more predictive than venue. F…

2024

Mitigating Translationese in Low-resource Languages: The Storyboard Approach

COLING 2024main

Low-resource languages often face challenges in acquiring high-quality language data due to the reliance on translation-based methods, which can introduce the translationese effect. This phenomenon results in translated sentences that lack fluency and naturalness in the target language. In this pape…

Cited by 2SourcePDFScholar
2024

No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages

EMNLP 2024main

Research in vision and language has made considerable progress thanks to benchmarks such as COCO. COCO captions focused on unambiguous facts in English; ArtEmis introduced subjective emotions and ArtELingo introduced some multilinguality (Chinese and Arabic). However we believe there should be more…

2023

AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages

EMNLP 2023long main

Africa is home to over 2,000 languages from over six language families and has the highest linguistic diversity among all continents. This includes 75 languages with at least one million speakers each. Yet, there is little NLP research conducted on African languages. Crucial in enabling such researc…

Cited by 0SourcecodeScholar
2023

Cross-lingual Open-Retrieval Question Answering for African Languages

EMNLP 2023long findings

African languages have far less in-language content available digitally, making it challenging for question answering systems to satisfy the information needs of users. Cross-lingual open-retrieval question answering (XOR QA) systems -- those that retrieve answer content from other languages while s…

Cited by 0SourceScholar
2023

HaVQA: A Dataset for Visual Question Answering and Multimodal Research in Hausa Language

ACL 2023findings

This paper presents “HaVQA”, the first multimodal dataset for visual question answering (VQA) tasks in the Hausa language. The dataset was created by manually translating 6,022 English question-answer pairs, which are associated with 1,555 unique images from the Visual Genome dataset. As a result, t…