← Search

Xiaojun Qi

2 accepted papers

2026

Broadening the Backdoor Basin: Understanding LLM Backdoors Collapse and Making Backdoors Persistent

ICML 2026poster

Large Language Models (LLMs) have shown to be vulnerable to backdoor attacks, yet we observe that many LLM backdoors do not survive when end users perform supervised fine-tuning (SFT). In this work, we provide a geometric explanation: by probing the backdoor objective under controlled weight perturb…

Cited by 0SourceScholar
2026

Don't Shift the Trigger: Robust Gradient Ascent for Backdoor Unlearning

ICLR 2026poster

Backdoor attacks pose a significant threat to machine learning models, allowing adversaries to implant hidden triggers that alter model behavior when activated. Although gradient ascent (GA)-based unlearning has been proposed as an efficient backdoor removal approach, we identify a critical yet over…

Cited by 0SourceScholar