← Search

Adaku Uchendu

10 accepted papers

2026

Position: Breaking the Dual Curse of Multilingual AI Requires Socio-Technical Guardrails, Not Post-Hoc Alignment

ICML 2026poster

Large language models are deployed globally as universal systems, yet their safety mechanisms remain English-optimized. This creates a Dual Curse for speakers of low-resource languages: a Harmfulness Curse where harmful content generation rises from 1\% in English to 35\% in languages like Hausa, Ig…

Cited by 0SourceScholar
2025

Beemo: Benchmark of Expert-edited Machine-generated Outputs

NAACL 2025long

The rapid proliferation of large language models (LLMs) has increased the volume of machine-generated texts (MGTs) and blurred text authorship in various domains. However, most existing MGT benchmarks include single-author texts (human-written and machine-generated). This conventional design fails t…

2025

PlagBench: Exploring the Duality of Large Language Models in Plagiarism Generation and Detection

NAACL 2025long

Recent studies have raised concerns about the potential threats large language models (LLMs) pose to academic integrity and copyright protection. Yet, their investigation is predominantly focused on literal copies of original texts. Also, how LLMs can facilitate the detection of LLM-generated plagia…

2024

A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts

ACL 2024long

In the realm of text manipulation and linguistic transformation, the question of authorship has been a subject of fascination and philosophical inquiry. Much like the Ship of Theseus paradox, which ponders whether a ship remains the same when each of its original planks is replaced, our research del…

2024

Authorship Obfuscation in Multilingual Machine-Generated Text Detection

EMNLP 2024finding

High-quality text generation capability of latest Large Language Models (LLMs) causes concerns about their misuse (e.g., in massive generation/spread of disinformation). Machine-generated text (MGT) detection is important to cope with such threats. However, it is susceptible to authorship obfuscatio…

2024

GPT-who: An Information Density-based Machine-Generated Text Detector

NAACL 2024findings

The Uniform Information Density (UID) principle posits that humans prefer to spread information evenly during language production. We examine if this UID principle can help capture differences between Large Language Models (LLMs)-generated and human-generated texts. We propose GPT-who, the first psy…

2023

Fighting Fire with Fire: The Dual Role of LLMs in Crafting and Detecting Elusive Disinformation

EMNLP 2023long main

Recent ubiquity and disruptive impacts of large language models (LLMs) have raised concerns about their potential to be misused (*.i.e, generating large-scale harmful and misleading content*). To combat this emerging risk of LLMs, we propose a novel "***Fighting Fire with Fire***" (F3) strategy that…

Cited by 0SourcecodeScholar
2023

HANSEN: Human and AI Spoken Text Benchmark for Authorship Analysis

EMNLP 2023long findings

$\textit{Authorship Analysis}$, also known as stylometry, has been an essential aspect of Natural Language Processing (NLP) for a long time. Likewise, the recent advancement of Large Language Models (LLMs) has made authorship analysis increasingly crucial for distinguishing between human-written and…

Cited by 0SourceScholar
2023

MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark

EMNLP 2023long main

There is a lack of research into capabilities of recent LLMs to generate convincing text in languages other than English and into performance of detectors of machine-generated text in multilingual settings. This is also reflected in the available benchmarks which lack authentic texts in languages ot…

Cited by 0SourcecodeScholar
2021

TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation

EMNLP 2021finding

Recent progress in generative language models has enabled machines to generate astonishingly realistic texts. While there are many legitimate applications of such models, there is also a rising need to distinguish machine-generated texts from human-written ones (e.g., fake news detection). However,…