2026
STAR: Strategy-driven Automatic Jailbreak Red-teaming For Large Language Model
ICLR 2026poster
Jailbreaking refers to techniques that bypass the safety alignment of large language models (LLMs) to elicit harmful outputs, and automated red-teaming has become a key approach for detecting such vulnerabilities before deployment. However, most existing red-teaming methods operate directly in text…