2026
Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks
AAAI 2026technical
Backdoor attacks pose a serious threat to the security of large language models (LLMs), causing them to exhibit anomalous behavior under specific trigger conditions. The design of backdoor triggers has evolved from fixed triggers to dynamic or implicit triggers. This increased flexibility in trigger