← Search

Anna Sztyber-Betley

2 accepted papers

2025

Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

ICML 2025oral

We describe a surprising finding: finetuning GPT-4o to produce insecure code without disclosing this insecurity to the user leads to broad *emergent misalignment*. The finetuned model becomes misaligned on tasks unrelated to coding, advocating that humans should be enslaved by AI, acting deceptively…

2025

Tell me about yourself: LLMs are aware of their learned behaviors

ICLR 2025spotlight

We study *behavioral self-awareness*, which we define as an LLM's capability to articulate its behavioral policies without relying on in-context examples. We finetune LLMs on examples that exhibit particular behaviors, including (a) making risk-seeking / risk-averse economic decisions, and (b) makin…