2026
Rethinking Deep Alignment Through the Lens of Incomplete Safety Learning
AAAI 2026technical
Large language models exhibit systematic vulnerabilities to adversarial attacks despite extensive safety alignment through supervised fine-tuning and reinforcement learning from human feedback. These vulnerabilities manifest as differential safety behavior across token positions, with safety modific