2025
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
ICML 2025spotlight
Methods for knowledge editing and unlearning in large language models seek to edit or remove undesirable knowledge or capabilities without compromising general language modeling performance. This work investigates how mechanistic interpretability---which, in part, aims to identify model components (…