2024
PKAD: Pretrained Knowledge is All You Need to Detect and Mitigate Textual Backdoor Attacks
EMNLP 2024finding
In textual backdoor attacks, attackers insert poisoned samples with triggered inputs and target labels into training datasets to manipulate model behavior, threatening the model’s security and reliability. Current defense methods can generally be categorized into inference-time and training-time one…