2026
Reflector: Internalizing Step-wise Reflection against Indirect Jailbreaks
ICML 2026poster
While Large Language Models (LLMs) demonstrate remarkable capabilities, they remain susceptible to sophisticated, multi-step jailbreak attacks that circumvent conventional surface-level safety alignment by exploiting the internal generation process. To address these vulnerabilities, we propose Refle…