AAAI 2026technical0 citations

Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning

Chenyu Zhang, Lanjun Wang, Yiwen Ma, Wenhui Li, Guoqing Jin, Anan Liu

Abstract

Text-to-Image (T2I) models typically deploy safety mechanisms to prevent the generation of sensitive images. Unfortunately, recent jailbreaking attack methods manually design instructions for the LLM to generate adversarial prompts, which effectively exposing safety vulnerabilities of T2I models. However, existing methods have two limitations: 1) relying on manually exhaustive strategies for designing adversarial prompts, lacking a unified framework, and 2) requiring numerous queries to achieve a successful attack, limiting their practical applicability. To address this issue, we propose Reason2Attack~(R2A), which aims to enhance the effectiveness and efficiency of the LLM in jailbreaking attacks. Specifically, we first use Frame Semantics theory to systematize existing manually crafted strategies and propose a unified generation framework to generate CoT adversarial prompts step by step. Following this, we propose a two-stage LLM reasoning training framework guided by the attack process. In the first stage, the LLM is fine-tuned with CoT examples generated by the unified generation framework to internalize the adversarial prompt generation process grounded in Frame Semantics. In the second stage, we incorporate the jailbreaking task into the LLM

BibTeX
@inproceedings{aaai2026_reason2attackjai,
  title = {Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning},
  author = {Chenyu Zhang and Lanjun Wang and Yiwen Ma and Wenhui Li and Guoqing Jin and Anan Liu},
  booktitle = {AAAI 2026},
  year = {2026}
}
Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning · AAAI 2026