2025
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
ICML 2025oral
We describe a surprising finding: finetuning GPT-4o to produce insecure code without disclosing this insecurity to the user leads to broad *emergent misalignment*. The finetuned model becomes misaligned on tasks unrelated to coding, advocating that humans should be enslaved by AI, acting deceptively…