← Search

Andy Arditi

2 accepted papers

2024

Refusal in Language Models Is Mediated by a Single Direction

NeurIPS 2024poster

Conversational large language models are fine-tuned for both instruction-following and safety, resulting in models that obey benign requests but refuse harmful ones. While this refusal behavior is widespread across chat models, its underlying mechanisms remain poorly understood. In this work, we sho…