← Search

Nina Rimsky

3 accepted papers

2024

Refusal in Language Models Is Mediated by a Single Direction

NeurIPS 2024poster

Conversational large language models are fine-tuned for both instruction-following and safety, resulting in models that obey benign requests but refuse harmful ones. While this refusal behavior is widespread across chat models, its underlying mechanisms remain poorly understood. In this work, we sho…

2024

Steering Llama 2 via Contrastive Activation Addition

ACL 2024long

We introduce Contrastive Activation Addition (CAA), a method for steering language models by modifying their activations during forward passes. CAA computes “steering vectors” by averaging the difference in residual stream activations between pairs of positive and negative examples of a particular b…