← Search

Hamdy Mubarak

10 accepted papers

2025

Advancing Arabic Diacritization: Improved Datasets, Benchmarking, and State-of-the-Art Models

EMNLP 2025

Arabic diacritics, similar to short vowels in English, provide phonetic and grammatical information but are typically omitted in written Arabic, leading to ambiguity. Diacritization (aka diacritic restoration or vowelization) is essential for natural language processing. This paper advances Arabic d

Cited by 0SourcePDFScholar
2024

Beyond Orthography: Automatic Recovery of Short Vowels and Dialectal Sounds in Arabic

ACL 2024long

This paper presents a novel Dialectal Sound and Vowelization Recovery framework, designed to recognize borrowed and dialectal sounds within phonologically diverse and dialect-rich languages, that extends beyond its standard orthographic sound sets. The proposed framework utilized quantized sequence…

2024

Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification

ACL 2024findings

Large language models (LLMs) are notorious for hallucinating, i.e., producing erroneous claims in their output. Such hallucinations can be dangerous, as occasional factual inaccuracies in the generated text might be obscured by the rest of the output being generally factually correct, making it extr…

2024

Halwasa: Quantify and Analyze Hallucinations in Large Language Models: Arabic as a Case Study

COLING 2024main

Large Language Models (LLMs) have shown superb abilities to generate texts that are indistinguishable from human-generated texts in many cases. However, sometimes they generate false, incorrect, or misleading content, which is often described as “hallucinations”. Quantifying and analyzing hallucinat…

Cited by 1SourcePDFScholar
2024

So Hateful! Building a Multi-Label Hate Speech Annotated Arabic Dataset

COLING 2024main

Social media enables widespread propagation of hate speech targeting groups based on ethnicity, religion, or other characteristics. With manual content moderation being infeasible given the volume, automatic hate speech detection is essential. This paper analyzes 70,000 Arabic tweets, from which 15,…

Cited by 4SourcePDFScholar
2021

Fighting the COVID-19 Infodemic: Modeling the Perspective of Journalists, Fact-Checkers, Social Media Platforms, Policy Makers, and the Society

EMNLP 2021finding

With the emergence of the COVID-19 pandemic, the political and the medical aspects of disinformation merged as the problem got elevated to a whole new level to become the first global infodemic. Fighting this infodemic has been declared one of the most important focus areas of the World Health Organ…

2021

QASR: QCRI Aljazeera Speech Resource A Large Scale Annotated Arabic Speech Corpus

ACL 2021long

We introduce the largest transcribed Arabic speech corpus, QASR, collected from the broadcast domain. This multi-dialect speech dataset contains 2,000 hours of speech sampled at 16kHz crawled from Aljazeera news channel. The dataset is released with lightly supervised transcriptions, aligned with th…

2020

ADI17: A Fine-Grained Arabic Dialect Identification Dataset

ICASSP 2020accepted

In this paper, we describe a method to collect dialectal speech from YouTube videos to create a large-scale Dialect Identification (DID) dataset. Using this method, we collected dialectal Arabic from known YouTube channels from 17 Arabic speaking countries in the Middle East and Northern Africa. Aft…

Cited by 0SourceScholar