2026
Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
ICLR 2026poster
Large Language Models (LLMs) exhibit strong but shallow alignment: they directly refuse harmful queries when a refusal is expected at the very start of an assistant turn, yet this protection collapses once a harmful continuation is underway (either through the adversarial attacks or via harmful assi…