AAAI 2026technical0 citations
ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
Zihan Wang, Rui Zhang, Hongwei Li, Wenshu Fan, Wenbo Jiang, Qingchuan Zhao, Guowen Xu
Abstract
Backdoor attacks pose a significant threat to Large Language Models (LLMs), where adversaries can embed hidden triggers to manipulate LLM
BibTeX
@inproceedings{aaai2026_confguardasimple,
title = {ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models},
author = {Zihan Wang and Rui Zhang and Hongwei Li and Wenshu Fan and Wenbo Jiang and Qingchuan Zhao and Guowen Xu},
booktitle = {AAAI 2026},
year = {2026}
}