← Search

Alexander Turner

3 accepted papers

2026

Recontextualization Mitigates Specification Gaming Without Modifying the Specification

ICML 2026poster

Developers often struggle to specify correct training labels and rewards. Perhaps they don't need to. We propose recontextualization, which reduces how often language models "game" training signals, performing misbehaviors those signals fail to penalize. We show recontextualization prevents models f…

Cited by 0SourceScholar
2024

Steering Llama 2 via Contrastive Activation Addition

ACL 2024long

We introduce Contrastive Activation Addition (CAA), a method for steering language models by modifying their activations during forward passes. CAA computes “steering vectors” by averaging the difference in residual stream activations between pairs of positive and negative examples of a particular b…

2019

Robustness May Be at Odds with Accuracy

ICLR 2019poster

We show that there exists an inherent tension between the goal of adversarial robustness and that of standard generalization. Specifically, training robust models may not only be more resource-consuming, but also lead to a reduction of standard accuracy. We demonstrate that this trade-off between t…

Cited by 2099SourcePDFScholar