← Search

Edward Turner

1 accepted papers

2026

Emergent Misalignment is Easy, Narrow Misalignment is Hard

ICLR 2026poster

Finetuning large language models on narrowly harmful datasets can cause them to become emergently misaligned, giving stereotypically `evil' responses across diverse unrelated settings. Concerningly, a pre-registered survey of experts failed to predict this result, highlighting our poor understanding…

Cited by 0SourcecodeScholar