← Search

Felix Hofstätter

3 accepted papers

2025

AI Sandbagging: Language Models can Strategically Underperform on Evaluations

ICLR 2025poster

Trustworthy capability evaluations are crucial for ensuring the safety of AI systems, and are becoming a key component of AI regulation. However, the developers of an AI system, or the AI system itself, may have incentives for evaluations to understate the AI's actual capability. These conflicting i…

2025

Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models

NeurIPS 2025poster

Capability evaluations play a crucial role in assessing and regulating frontier AI systems. The effectiveness of these evaluations faces a significant challenge: strategic underperformance, or ``sandbagging'', where models deliberately underperform during evaluation. Sandbagging can manifest either…

Cited by 0SourcecodeScholar
2025

The Elicitation Game: Evaluating Capability Elicitation Techniques

ICML 2025poster

Capability evaluations are required to understand and regulate AI systems that may be deployed or further developed. Therefore, it is important that evaluations provide an accurate estimation of an AI system’s capabilities. However, in numerous cases, previously latent capabilities have been elicite…

Cited by 0SourcePDFScholar