2026
SpatialJB: How Text Distribution Art Becomes The "Jailbreak Key" for LLM Guardrails
ICML 2026poster
While Large Language Models (LLMs) have achieved remarkable success across diverse tasks, they remain vulnerable to jailbreak attacks, which pose significant risks to their secure deployment. Current safetymechanisms primarily rely on output guardrails to filter harmful outputs, yet these defenses a…