← Search

Andrew Qin

1 accepted papers

2026

Steering Evaluation-Aware Language Models To Act Like They Are Deployed

ICLR 2026poster

Large language models (LLMs) can sometimes detect when they are being evaluated and adjust their behavior to appear more aligned, compromising the reliability of safety evaluations. In this paper, we show that adding a steering vector to an LLM's activations can suppress evaluation-awareness and mak…

Cited by 0SourcecodeScholar