← Search

Tim Tian Hua

2 accepted papers

2026

Steering Evaluation-Aware Language Models To Act Like They Are Deployed

ICLR 2026poster

Large language models (LLMs) can sometimes detect when they are being evaluated and adjust their behavior to appear more aligned, compromising the reliability of safety evaluations. In this paper, we show that adding a steering vector to an LLM's activations can suppress evaluation-awareness and mak…

Cited by 0SourcecodeScholar
2025

Combining Cost Constrained Runtime Monitors for AI Safety

NeurIPS 2025poster

Monitoring AIs at runtime can help us detect and stop harmful actions. In this paper, we study how to efficiently combine multiple runtime monitors into a single monitoring protocol. The protocol's objective is to maximize the probability of applying a safety intervention on misaligned outputs (i.e.…

Cited by 0SourceScholar