2026
Mitigating Safety Fallback in Editing-based Backdoor Injection on LLMs
ICLR 2026poster
Large language models (LLMs) have shown strong performance across natural language tasks, but remain vulnerable to backdoor attacks. Recent model editing-based approaches enable efficient backdoor injection by directly modifying parameters to map specific triggers to attacker-desired responses. Howe…