2026
Removing Sandbagging in LLMs by Training with Weak Supervision
ICML 2026poster
As AI systems begin to automate complex tasks, supervision increasingly relies on weaker models or limited human oversight that cannot fully verify output quality. A model more capable than its supervisors could exploit this gap through sandbagging, producing work that appears acceptable but falls s…