← Search

Thong Bach

1 accepted papers

2026

Rethinking Deep Alignment Through the Lens of Incomplete Safety Learning

AAAI 2026technical

Large language models exhibit systematic vulnerabilities to adversarial attacks despite extensive safety alignment through supervised fine-tuning and reinforcement learning from human feedback. These vulnerabilities manifest as differential safety behavior across token positions, with safety modific

Cited by 0SourcePDFScholar