2026
Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning
AAAI 2026technical
Text-to-Image (T2I) models typically deploy safety mechanisms to prevent the generation of sensitive images. Unfortunately, recent jailbreaking attack methods manually design instructions for the LLM to generate adversarial prompts, which effectively exposing safety vulnerabilities of T2I models.