AAAI 2026technical0 citations

ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models

Zihan Wang, Rui Zhang, Hongwei Li, Wenshu Fan, Wenbo Jiang, Qingchuan Zhao, Guowen Xu

Abstract

Backdoor attacks pose a significant threat to Large Language Models (LLMs), where adversaries can embed hidden triggers to manipulate LLM

BibTeX
@inproceedings{aaai2026_confguardasimple,
  title = {ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models},
  author = {Zihan Wang and Rui Zhang and Hongwei Li and Wenshu Fan and Wenbo Jiang and Qingchuan Zhao and Guowen Xu},
  booktitle = {AAAI 2026},
  year = {2026}
}
ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models · AAAI 2026