← Search

Marco Christiani

1 accepted papers

2025

Concept-ROT: Poisoning Concepts in Large Language Models with Model Editing

ICLR 2025poster

Model editing methods modify specific behaviors of Large Language Models by altering a small, targeted set of network weights and require very little data and compute. These methods can be used for malicious applications such as inserting misinformation or simple trojans that result in adversary-spe…