ICLR 2026poster0 citations

Persona Features Control Emergent Misalignment

Miles Wang, Tom Dupre la Tour, Olivia Watkins, Aleksandar Makelov, Ryan Andrew Chi, Samuel Miserendino, Jeffrey George Wang, Achyuta Rajaram

Abstract

Understanding how language models generalize behaviors from their training to a broader deployment distribution is an important problem in AI safety. Betley et al. discovered that fine-tuning GPT-4o on intentionally insecure code causes "emergent misalignment," where models give stereotypically malicious responses to unrelated prompts. We extend this work, demonstrating emergent misalignment across diverse conditions, including reinforcement learning on reasoning models, fine-tuning on various synthetic datasets, and in models without safety training. To investigate the mechanisms behind this generalized misalignment, we apply a "model diffing" approach using sparse autoencoders to compare internal model representations before and after fine-tuning. This approach reveals several "misaligned persona" features in activation space, including a toxic persona feature which most strongly controls emergent misalignment and can be used to predict whether a model will exhibit such behavior. Additionally, we investigate mitigation strategies, discovering that fine-tuning an emergently misaligned model on just a few hundred benign samples efficiently restores alignment.

interpretabilityalignmentsafety
BibTeX
@inproceedings{
wang2026persona,
title={Persona Features Control Emergent Misalignment},
author={Miles Wang and Tom Dupre la Tour and Olivia Watkins and Aleksandar Makelov and Ryan Andrew Chi and Samuel Miserendino and Jeffrey George Wang and Achyuta Rajaram and Johannes Heidecke and Tejal Patwardhan and Daniel P Mossing},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=yjrVOxjkDR}
}
Persona Features Control Emergent Misalignment · ICLR 2026