← Search

Nohayr Muhammad Abdelmoneim

2 accepted papers

2025

Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs

EMNLP 2025

In this work, we introduce SPLICE, a human-curated benchmark derived from the COIN instructional video dataset, designed to probe event-based reasoning across multiple dimensions: temporal, causal, spatial, contextual, and general knowledge. SPLICE includes 3,381 human-filtered videos spanning 12 ca

Cited by 0SourcePDFScholar
2024

FOOL ME IF YOU CAN! An Adversarial Dataset to Investigate the Robustness of LMs in Word Sense Disambiguation

EMNLP 2024main

Word sense disambiguation (WSD) is a key task in natural language processing and lexical semantics. Pre-trained language models with contextualized word embeddings have significantly improved performance in regular WSD tasks. However, these models still struggle with recognizing semantic boundaries…

Cited by 1SourcePDFScholar