2025
Neuroplasticity and Corruption in Model Mechanisms: A Case Study Of Indirect Object Identification
NAACL 2025findings
Previous research has shown that fine-tuning language models on general tasks enhance their underlying mechanisms. However, the impact of fine-tuning on poisoned data and the resulting changes in these mechanisms are poorly understood. This study investigates the changes in a model’s mechanisms duri…