2026
Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM Judges
ICLR 2026poster
Identifying the vulnerabilities of large language models (LLMs) is crucial for improving their safety by addressing inherent weaknesses. Jailbreaks, in which adversaries bypass safeguards with crafted input prompts, play a central role in red-teaming by probing LLMs to elicit unintended or unsafe be…