ACL 2023findings290 citations

Discovering Language Model Behaviors with Model-Written Evaluations

Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson

Abstract

As language models (LMs) scale, they develop many novel behaviors, good and bad, exacerbating the need to evaluate how they behave. Prior work creates evaluations with crowdwork (which is time-consuming and expensive) or existing data sources (which are not always available). Here, we automatically generate evaluations with LMs. We explore approaches with varying amounts of human effort, from instructing LMs to write yes/no questions to making complex Winogender schemas with multiple stages of LM-based generation and filtering. Crowdworkers rate the examples as highly relevant and agree with 90-100% of labels, sometimes more so than corresponding human-written datasets. We generate 154 datasets and discover new cases of inverse scaling where LMs get worse with size. Larger LMs repeat back a dialog user’s preferred answer (“sycophancy”) and express greater desire to pursue concerning goals like resource acquisition and goal preservation. We also find some of the first examples of inverse scaling in RL from Human Feedback (RLHF), where more RLHF makes LMs worse. For example, RLHF makes LMs express stronger political views (on gun rights and immigration) and a greater desire to avoid shut down. Overall, LM-written evaluations are high-quality and let us quickly discover many novel LM behaviors.

BibTeX
@inproceedings{perez-etal-2023-discovering,
    title = "Discovering Language Model Behaviors with Model-Written Evaluations",
    author = "Perez, Ethan  and
      Ringer, Sam  and
      Lukosiute, Kamile  and
      Nguyen, Karina  and
      Chen, Edwin  and
      Heiner, Scott  and
      Pettit, Craig  and
      Olsson, Catherine  and
      Kundu, Sandipan  and
      Kadavath, Saurav  and
      Jones, Andy  and
      Chen, Anna  and
      Mann, Benjamin  and
      Israel, Brian  and
      Seethor, Bryan  and
      McKinnon, Cameron  and
      Olah, Christopher  and
      Yan, Da  and
      Amodei, Daniela  and
      Amodei, Dario  and
      Drain, Dawn  and
      Li, Dustin  and
      Tran-Johnson, Eli  and
      Khundadze, Guro  and
      Kernion, Jackson  and
      Landis, James  and
      Kerr, Jamie  and
      Mueller, Jared  and
      Hyun, Jeeyoon  and
      Landau, Joshua  and
      Ndousse, Kamal  and
      Goldberg, Landon  and
      Lovitt, Liane  and
      Lucas, Martin  and
      Sellitto, Michael  and
      Zhang, Miranda  and
      Kingsland, Neerav  and
      Elhage, Nelson  and
      Joseph, Nicholas  and
      Mercado, Noemi  and
      DasSarma, Nova  and
      Rausch, Oliver  and
      Larson, Robin  and
      McCandlish, Sam  and
      Johnston, Scott  and
      Kravec, Shauna  and
      El Showk, Sheer  and
      Lanham, Tamera  and
      Telleen-Lawton, Timothy  and
      Brown, Tom  and
      Henighan, Tom  and
      Hume, Tristan  and
      Bai, Yuntao  and
      Hatfield-Dodds, Zac  and
      Clark, Jack  and
      Bowman, Samuel R.  and
      Askell, Amanda  and
      Grosse, Roger  and
      Hernandez, Danny  and
      Ganguli, Deep  and
      Hubinger, Evan  and
      Schiefer, Nicholas  and
      Kaplan, Jared",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2023",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.findings-acl.847/",
    doi = "10.18653/v1/2023.findings-acl.847",
    pages = "13387--13434"
}
Discovering Language Model Behaviors with Model-Written Evaluations · ACL 2023