2024
Who's asking? User personas and the mechanics of latent misalignment
NeurIPS 2024spotlight
Studies show that safety-tuned models may nevertheless divulge harmful information. In this work, we show that whether they do so depends significantly on who they are talking to, which we refer to as *user persona*. In fact, we find manipulating user persona to be more effective for eliciting harmf…