← Search

Weiliang Zhao

2 accepted papers

2025

Diversity Helps Jailbreak Large Language Models

NAACL 2025long

We have uncovered a powerful jailbreak technique that leverages large language models’ ability to diverge from prior context, enabling them to bypass safety constraints and generate harmful outputs. By simply instructing the LLM to deviate and obfuscate previous attacks, our method dramatically outp…

Cited by 2SourcePDFScholar
2025

Learning to Rewrite: Generalized LLM-Generated Text Detection

ACL 2025long

Detecting text generated by Large Language Models (LLMs) is crucial, yet current detectors often struggle to generalize in open-world settings. We introduce Learning2Rewrite, a novel framework to detect LLM-generated text with exceptional generalization to unseen domains. Capitalized on the finding…