2026
Steering Evaluation-Aware Language Models To Act Like They Are Deployed
ICLR 2026poster
Large language models (LLMs) can sometimes detect when they are being evaluated and adjust their behavior to appear more aligned, compromising the reliability of safety evaluations. In this paper, we show that adding a steering vector to an LLM's activations can suppress evaluation-awareness and mak…