← Search

Niels Warncke

3 accepted papers

2026

Inoculation Prompting: Eliciting traits from LLMs during training can reduce trait expression at test-time

ICLR 2026poster

Language model finetuning often results in learning undesirable traits in combination with desired ones. To address this, we propose inoculation prompting: modifying finetuning data by prepending a short system-prompt instruction that deliberately elicits the undesirable trait. At test time, we eval…

Cited by 0SourceScholar
2025

Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

ICML 2025oral

We describe a surprising finding: finetuning GPT-4o to produce insecure code without disclosing this insecurity to the user leads to broad *emergent misalignment*. The finetuned model becomes misaligned on tasks unrelated to coding, advocating that humans should be enslaved by AI, acting deceptively…