← Search

Aaquib Syed

2 accepted papers

2025

Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization

ICML 2025spotlight

Methods for knowledge editing and unlearning in large language models seek to edit or remove undesirable knowledge or capabilities without compromising general language modeling performance. This work investigates how mechanistic interpretability---which, in part, aims to identify model components (…

Cited by 10SourcePDFScholar
2024

Refusal in Language Models Is Mediated by a Single Direction

NeurIPS 2024poster

Conversational large language models are fine-tuned for both instruction-following and safety, resulting in models that obey benign requests but refuse harmful ones. While this refusal behavior is widespread across chat models, its underlying mechanisms remain poorly understood. In this work, we sho…