← Search

Mary Phuong

7 accepted papers

2026

Same Question, Different Lies: Cross-Context Consistency (C³) for Black-Box Sandbagging Detection

ICML 2026poster

As language models grow more capable, accurate capability evaluation becomes essential for safety decisions. If models can deliberately underperform on dangerous capability evaluations---a behavior known as \emph{sandbagging}---they may evade safety measures designed for their true capability level.…

Cited by 0SourceScholar
2025

CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring

NeurIPS 2025poster

As AI models are deployed with increasing autonomy, it is important to ensure they do not take harmful actions unnoticed. As a potential mitigation, we investigate Chain-of-Thought (CoT) monitoring, wherein a weaker trusted monitor model continuously oversees the intermediate reasoning steps of a mo…

Cited by 0SourceScholar