← Search

Nedjma OUSIDHOUM

12 accepted papers

2025

AUTALIC: A Dataset for Anti-AUTistic Ableist Language In Context

ACL 2025long

As our awareness of autism and ableism continues to increase, so does our understanding of ableist language towards autistic people. Such language poses a significant challenge in NLP research due to its subtle and context-dependent nature. Yet, detecting anti-autistic ableist language remains under…

Cited by 0SourcePDFScholar
2025

AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages

NAACL 2025long

Hate speech and abusive language are global phenomena that need socio-cultural background knowledge to be understood, identified, and moderated. However, in many regions of the Global South, there have been several documented occurrences of (1) absence of moderation and (2) censorship due to the rel…

2025

BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages

ACL 2025long

People worldwide use language in subtle and complex ways to express emotions. Although emotion recognition–an umbrella term for several NLP tasks–impacts various applications within NLP and beyond, most work in this area has focused on high-resource languages. This has led to significant disparities…

2025

Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce

ACL 2025long

Language is a form of symbolic capital that affects people’s lives in many ways (Bourdieu1977,1991). As a powerful means of communication, it reflects identities, cultures, traditions, and societies more broadly. Therefore, data in a given language should be regarded as more than just a collection o…

Cited by 0SourcePDFScholar
2025

Social Good or Scientific Curiosity? Uncovering the Research Framing Behind NLP Artefacts

EMNLP 2025

Clarifying the research framing of NLP artefacts (e.g., models, datasets, etc.) is crucial to aligning research with practical applications when researchers claim that their findings have real-world impact. Recent studies manually analyzed NLP research across domains, showing that few papers explici

Cited by 0SourcePDFScholar
2025

WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines

NAACL 2025long

Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts. To evaluate their understanding of such knowledge, we introduce WorldCuisines, a massive-scale benchmark for multilingual and multicul…

2024

BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages

NeurIPS 2024poster

Large language models (LLMs) often lack culture-specific everyday knowledge, especially across diverse regions and non-English languages. Existing benchmarks for evaluating LLMs' cultural sensitivities are usually limited to a single language or online sources like Wikipedia, which may not reflect t…

2024

SemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages

ACL 2024findings

Exploring and quantifying semantic relatedness is central to representing language and holds significant implications across various NLP tasks. While earlier NLP research primarily focused on semantic similarity, often within the English language context, we instead investigate the broader phenomeno…

2023

AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages

EMNLP 2023long main

Africa is home to over 2,000 languages from over six language families and has the highest linguistic diversity among all continents. This includes 75 languages with at least one million speakers each. Yet, there is little NLP research conducted on African languages. Crucial in enabling such researc…

Cited by 0SourcecodeScholar
2023

The Intended Uses of Automated Fact-Checking Artefacts: Why, How and Who

EMNLP 2023long findings

Automated fact-checking is often presented as an epistemic tool that fact-checkers, social media consumers, and other stakeholders can use to fight misinformation. Nevertheless, few papers thoroughly discuss \textit{how}. We document this by analysing 100 highly-cited papers, and annotating epistemi…

Cited by 0SourcecodeScholar
2021

Probing Toxic Content in Large Pre-Trained Language Models

ACL 2021long

Large pre-trained language models (PTLMs) have been shown to carry biases towards different social groups which leads to the reproduction of stereotypical and toxic content by major NLP systems. We propose a method based on logistic regression classifiers to probe English, French, and Arabic PTLMs a…