← Search

Aidan Ewart

2 accepted papers

2025

Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization

ICML 2025spotlight

Methods for knowledge editing and unlearning in large language models seek to edit or remove undesirable knowledge or capabilities without compromising general language modeling performance. This work investigates how mechanistic interpretability---which, in part, aims to identify model components (…

Cited by 10SourcePDFScholar
2024

Sparse Autoencoders Find Highly Interpretable Features in Language Models

ICLR 2024poster

One of the roadblocks to a better understanding of neural networks' internals is \textit{polysemanticity}, where neurons appear to activate in multiple, semantically distinct contexts. Polysemanticity prevents us from identifying concise, human-understandable explanations for what neural networks ar…