← Search

Maged S. Al-shaibani

3 accepted papers

2025

MOLE: Metadata Extraction and Validation in Scientific Papers Using LLMs

EMNLP 2025

Metadata extraction is essential for cataloging and preserving datasets, enabling effective research discovery and reproducibility, especially given the current exponential growth in scientific research. While Masader (CITATION) laid the groundwork for extracting a wide range of metadata attributes

2025

Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset

EMNLP 2025

Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets. To address this gap, we introduce PEARL, a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding. Constructed through

2023

Consonant is all you need: a compact representation of English text for efficient NLP

EMNLP 2023long findings

In natural language processing (NLP), the representation of text plays a crucial role in various tasks such as language modeling, sentiment analysis, and machine translation. The standard approach is to represent text in the same way as we, as humans, read and write. In this paper, we propose a nove…

Cited by 0SourceScholar