← Search

Emil Ryd

1 accepted papers

2026

Removing Sandbagging in LLMs by Training with Weak Supervision

ICML 2026poster

As AI systems begin to automate complex tasks, supervision increasingly relies on weaker models or limited human oversight that cannot fully verify output quality. A model more capable than its supervisors could exploit this gap through sandbagging, producing work that appears acceptable but falls s…

Cited by 0SourceScholar