2025
When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations
ACL 2025long
Large Language Models (LLMs) are known to be vulnerable to backdoor attacks, where triggers embedded in poisoned samples can maliciously alter LLMs’ behaviors. In this paper, we move beyond attacking LLMs and instead examine backdoor attacks through the novel lens of natural language explanations. S…