← Search

Daniel Ben-Levi

1 accepted papers

2025

Diversity Helps Jailbreak Large Language Models

NAACL 2025long

We have uncovered a powerful jailbreak technique that leverages large language models’ ability to diverge from prior context, enabling them to bypass safety constraints and generate harmful outputs. By simply instructing the LLM to deviate and obfuscate previous attacks, our method dramatically outp…

Cited by 2SourcePDFScholar