← Search

Henning Bartsch

2 accepted papers

2026

Removing Sandbagging in LLMs by Training with Weak Supervision

ICML 2026poster

As AI systems begin to automate complex tasks, supervision increasingly relies on weaker models or limited human oversight that cannot fully verify output quality. A model more capable than its supervisors could exploit this gap through sandbagging, producing work that appears acceptable but falls s…

Cited by 1SourceScholar
2025

The Elicitation Game: Evaluating Capability Elicitation Techniques

ICML 2025poster

Capability evaluations are required to understand and regulate AI systems that may be deployed or further developed. Therefore, it is important that evaluations provide an accurate estimation of an AI system’s capabilities. However, in numerous cases, previously latent capabilities have been elicite…

Cited by 0SourcePDFScholar