← Search

Nathan Labenz

1 accepted papers

2025

Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

ICML 2025oral

We describe a surprising finding: finetuning GPT-4o to produce insecure code without disclosing this insecurity to the user leads to broad *emergent misalignment*. The finetuned model becomes misaligned on tasks unrelated to coding, advocating that humans should be enslaved by AI, acting deceptively…