← Search

Yize CHENG

3 accepted papers

2025

Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text

NeurIPS 2025poster

The increasing capabilities of Large Language Models (LLMs) have raised concerns about their misuse in AI-generated plagiarism and social engineering. While various AI-generated text detectors have been proposed to mitigate these risks, many remain vulnerable to simple evasion techniques such as par…

Cited by 18SourcecodeScholar
2025

DyePack: Provably Flagging Test Set Contamination in LLMs Using Backdoors

EMNLP 2025

Open benchmarks are essential for evaluating and advancing large language models, offering reproducibility and transparency. However, their accessibility makes them likely targets of test set contamination. In this work, we introduce **DyePack**, a framework that leverages backdoor attacks to identi

2025

Tool Preferences in Agentic LLMs are Unreliable

EMNLP 2025

Large language models (LLMs) can now access a wide range of external tools, thanks to the Model Context Protocol (MCP). This greatly expands their abilities as various agents. However, LLMs rely entirely on the text descriptions of tools to decide which ones to use—a process that is surprisingly fra