← Search

Huaizhi Ge

1 accepted papers

2025

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations

ACL 2025long

Large Language Models (LLMs) are known to be vulnerable to backdoor attacks, where triggers embedded in poisoned samples can maliciously alter LLMs’ behaviors. In this paper, we move beyond attacking LLMs and instead examine backdoor attacks through the novel lens of natural language explanations. S…

Cited by 0SourcePDFScholar