AAAI 2026technical0 citations

Distractor-Based Jailbreaking Attacks in Language Models and Associated Changes in Chain-of-Thought Content (Student Abstract)

Tate Rowney, Xuning Ying

Abstract

We identify a jailbreaking vulnerability in multiple open-source LLMs: by augmenting dangerous requests using certain "distractors" to obfuscate their intent, we elicit specific, actionable responses on a wide variety of harmful topics. We find that such an attack noticeably alters the contents of these models

BibTeX
@inproceedings{aaai2026_distractorbasedj,
  title = {Distractor-Based Jailbreaking Attacks in Language Models and Associated Changes in Chain-of-Thought Content (Student Abstract)},
  author = {Tate Rowney and Xuning Ying},
  booktitle = {AAAI 2026},
  year = {2026}
}
Distractor-Based Jailbreaking Attacks in Language Models and Associated Changes in Chain-of-Thought Content (Student Abstract) · AAAI 2026