← Search

Nizar Habash

24 accepted papers

2026

MedAraBench: Large-scale Arabic Medical Question Answering Dataset and Benchmark

ICLR 2026poster

Arabic remains one of the most underrepresented languages in natural language processing research, particularly in medical applications, due to the limited availability of open-source data and benchmarks. The lack of resources hinders efforts to evaluate and advance the multilingual capabilities of…

Cited by 0SourcecodeScholar
2025

A Large and Balanced Corpus for Fine-grained Arabic Readability Assessment

ACL 2025finding

This paper introduces the Balanced Arabic Readability Evaluation Corpus (BAREC), a large-scale, fine-grained dataset for Arabic readability assessment. BAREC consists of 69,441 sentences spanning 1+ million words, carefully curated to cover 19 readability levels, from kindergarten to postgraduate co…

Cited by 0SourcePDFScholar
2025

A Survey of Code-switched Arabic NLP: Progress, Challenges, and Future Directions

COLING 2025main

Language in the Arab world presents a complex diglossic and multilingual setting, involving the use of Modern Standard Arabic, various dialects and sub-dialects, as well as multiple European languages. This diverse linguistic landscape has given rise to code-switching, both within Arabic varieties a…

Cited by 1SourcePDFScholar
2025

Data Augmentation for Maltese NLP using Transliterated and Machine Translated Arabic Data

EMNLP 2025

Maltese is a unique Semitic language that has evolved under extensive influence from Romance and Germanic languages, particularly Italian and English. Despite its Semitic roots, its orthography is based on the Latin script, creating a gap between it and its closest linguistic relatives in Arabic. In

2025

From Multiple-Choice to Extractive QA: A Case Study for English and Arabic

COLING 2025main

The rapid evolution of Natural Language Processing (NLP) has favoured major languages such as English, leaving a significant gap for many others due to limited resources. This is especially evident in the context of data annotation, a task whose importance cannot be underestimated, but which is time…

2024

Arabic Diacritics in the Wild: Exploiting Opportunities for Improved Diacritization

ACL 2024long

The widespread absence of diacritical marks in Arabic text poses a significant challenge for Arabic natural language processing (NLP). This paper explores instances of naturally occurring diacritics, referred to as “diacritics in the wild,” to unveil patterns and latent information across six divers…

2024

ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic

ACL 2024findings

The focus of language model evaluation has transitioned towards reasoning and knowledge-intensive tasks, driven by advancements in pretraining large models. While state-of-the-art models are partially trained on large Arabic texts, evaluating their performance in Arabic remains challenging due to th…

2024

Camel Morph MSA: A Large-Scale Open-Source Morphological Analyzer for Modern Standard Arabic

COLING 2024main

We present Camel Morph MSA, the largest open-source Modern Standard Arabic morphological analyzer and generator. Camel Morph MSA has over 100K lemmas, and includes rarely modeled morphological features of Modern Standard Arabic with Classical Arabic origins. Camel Morph MSA can produce ∼1.45B analys…

Cited by 2SourcePDFScholar
2024

LLM-DetectAIve: a Tool for Fine-Grained Machine-Generated Text Detection

EMNLP 2024system demonstrations

The ease of access to large language models (LLMs) has enabled a widespread of machine-generated texts, and now it is often hard to tell whether a piece of text was human-written or machine-generated. This raises concerns about potential misuse, particularly within educational and academic domains.…

2024

M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection

ACL 2024long

The advent of Large Language Models (LLMs) has brought an unprecedented surge in machine-generated text (MGT) across diverse channels. This raises legitimate concerns about its potential misuse and societal implications. The need to identify and differentiate such content from genuine human-generate…

2024

Palmyra 3.0: A User-Friendly Cloud-Based Platform for Morphology and Dependency Syntax Annotation

COLING 2024main

We present Palmyra 3.0, a cloud-based, configurable, and user-friendly platform for morphology and syntax annotation through dependency-tree visualization. Palmyra 3.0 implements a robust system that stores data on the cloud. By default, Palmyra 3.0 comes with an Arabic dependency parser that genera…

2024

The SAMER Arabic Text Simplification Corpus

COLING 2024main

We present the SAMER Corpus, the first manually annotated Arabic parallel corpus for text simplification targeting school-aged learners. Our corpus comprises texts of 159K words selected from 15 publicly available Arabic fiction novels most of which were published between 1865 and 1955. Our corpus i…

Cited by 7SourcePDFScholar
2024

ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus

COLING 2024main

We present ZAEBUC-Spoken, a multilingual multidialectal Arabic-English speech corpus. The corpus comprises twelve hours of Zoom meetings involving multiple speakers role-playing a work situation where Students brainstorm ideas for a certain topic and then discuss it with an Interlocutor. The meeting…

Cited by 6SourcePDFScholar
2023

Advancements in Arabic Grammatical Error Detection and Correction: An Empirical Investigation

EMNLP 2023long main

Grammatical error correction (GEC) is a well-explored problem in English with many existing models and datasets. However, research on GEC in morphologically rich languages has been limited due to challenges such as data scarcity and language complexity. In this paper, we present the first results on…

Cited by 0SourcecodeScholar
2022

Morphosyntactic Tagging with Pre-trained Language Models for Arabic and its Dialects

ACL 2022findings

We present state-of-the-art results on morphosyntactic tagging across different varieties of Arabic using fine-tuned pre-trained transformer language models. Our models consistently outperform existing systems in Modern Standard Arabic and all the Arabic dialects we study, achieving 2.6% absolute im…

2020

An Online Readability Leveled Arabic Thesaurus

COLING 2020system demonstrations

This demo paper introduces the online Readability Leveled Arabic Thesaurus interface. For a given user input word, this interface provides the word’s possible lemmas, roots, English glosses, related Arabic words and phrases, and readability on a five-level readability scale. This interface builds on…

Cited by 5SourcePDFScholar
2020

Multitask Easy-First Dependency Parsing: Exploiting Complementarities of Different Dependency Representations

COLING 2020main

In this paper we present a parsing model for projective dependency trees which takes advantage of the existence of complementary dependency annotations which is the case in Arabic, with the availability of CATiB and UD treebanks. Our system performs syntactic parsing according to both annotation typ…

2020

Utilizing Subword Entities in Character-Level Sequence-to-Sequence Lemmatization Models

COLING 2020main

In this paper we present a character-level sequence-to-sequence lemmatization model, utilizing several subword features in multiple configurations. In addition to generic n-gram embeddings (using FastText), we experiment with concatenative (stems) and templatic (roots and patterns) morphological sub…

Cited by 3SourcePDFScholar