← Search

Pablo Bernabeu-Perez

3 accepted papers

2026

Same Question, Different Lies: Cross-Context Consistency (C³) for Black-Box Sandbagging Detection

ICML 2026poster

As language models grow more capable, accurate capability evaluation becomes essential for safety decisions. If models can deliberately underperform on dangerous capability evaluations---a behavior known as \emph{sandbagging}---they may evade safety measures designed for their true capability level.…

Cited by 0SourceScholar
2025

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering

NAACL 2025short

Current Large Language Models (LLMs) benchmarks are often based on open-ended or close-ended QA evaluations, avoiding the requirement of human labor. Close-ended measurements evaluate the factuality of responses but lack expressiveness. Open-ended capture the model’s capacity to produce discourse re…

Cited by 1SourcePDFScholar
2025

CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring

NeurIPS 2025poster

As AI models are deployed with increasing autonomy, it is important to ensure they do not take harmful actions unnoticed. As a potential mitigation, we investigate Chain-of-Thought (CoT) monitoring, wherein a weaker trusted monitor model continuously oversees the intermediate reasoning steps of a mo…

Cited by 0SourceScholar