Broadening the Backdoor Basin: Understanding LLM Backdoors Collapse and Making Backdoors Persistent
Large Language Models (LLMs) have shown to be vulnerable to backdoor attacks, yet we observe that many LLM backdoors do not survive when end users perform supervised fine-tuning (SFT). In this work, we provide a geometric explanation: by probing the backdoor objective under controlled weight perturb…