2025
Detecting Stealthy Backdoor Samples based on Intra-class Distance for Large Language Models
EMNLP 2025
Stealthy data poisoning during fine-tuning can backdoor large language models (LLMs), threatening downstream safety. Existing detectors either use classifier-style probability signals—ill-suited to generation—or rely on rewriting, which can degrade quality and even introduce new triggers. We address