2025
Concept-ROT: Poisoning Concepts in Large Language Models with Model Editing
ICLR 2025poster
Model editing methods modify specific behaviors of Large Language Models by altering a small, targeted set of network weights and require very little data and compute. These methods can be used for malicious applications such as inserting misinformation or simple trojans that result in adversary-spe…