← Search

Julian Schulz

1 accepted papers

2024

Steering Llama 2 via Contrastive Activation Addition

ACL 2024long

We introduce Contrastive Activation Addition (CAA), a method for steering language models by modifying their activations during forward passes. CAA computes “steering vectors” by averaging the difference in residual stream activations between pairs of positive and negative examples of a particular b…