Discovering Language Model Behaviors with Model-Written Evaluations
Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson
Abstract
As language models (LMs) scale, they develop many novel behaviors, good and bad, exacerbating the need to evaluate how they behave. Prior work creates evaluations with crowdwork (which is time-consuming and expensive) or existing data sources (which are not always available). Here, we automatically generate evaluations with LMs. We explore approaches with varying amounts of human effort, from instructing LMs to write yes/no questions to making complex Winogender schemas with multiple stages of LM-based generation and filtering. Crowdworkers rate the examples as highly relevant and agree with 90-100% of labels, sometimes more so than corresponding human-written datasets. We generate 154 datasets and discover new cases of inverse scaling where LMs get worse with size. Larger LMs repeat back a dialog user’s preferred answer (“sycophancy”) and express greater desire to pursue concerning goals like resource acquisition and goal preservation. We also find some of the first examples of inverse scaling in RL from Human Feedback (RLHF), where more RLHF makes LMs worse. For example, RLHF makes LMs express stronger political views (on gun rights and immigration) and a greater desire to avoid shut down. Overall, LM-written evaluations are high-quality and let us quickly discover many novel LM behaviors.
BibTeX
@inproceedings{perez-etal-2023-discovering,
title = "Discovering Language Model Behaviors with Model-Written Evaluations",
author = "Perez, Ethan and
Ringer, Sam and
Lukosiute, Kamile and
Nguyen, Karina and
Chen, Edwin and
Heiner, Scott and
Pettit, Craig and
Olsson, Catherine and
Kundu, Sandipan and
Kadavath, Saurav and
Jones, Andy and
Chen, Anna and
Mann, Benjamin and
Israel, Brian and
Seethor, Bryan and
McKinnon, Cameron and
Olah, Christopher and
Yan, Da and
Amodei, Daniela and
Amodei, Dario and
Drain, Dawn and
Li, Dustin and
Tran-Johnson, Eli and
Khundadze, Guro and
Kernion, Jackson and
Landis, James and
Kerr, Jamie and
Mueller, Jared and
Hyun, Jeeyoon and
Landau, Joshua and
Ndousse, Kamal and
Goldberg, Landon and
Lovitt, Liane and
Lucas, Martin and
Sellitto, Michael and
Zhang, Miranda and
Kingsland, Neerav and
Elhage, Nelson and
Joseph, Nicholas and
Mercado, Noemi and
DasSarma, Nova and
Rausch, Oliver and
Larson, Robin and
McCandlish, Sam and
Johnston, Scott and
Kravec, Shauna and
El Showk, Sheer and
Lanham, Tamera and
Telleen-Lawton, Timothy and
Brown, Tom and
Henighan, Tom and
Hume, Tristan and
Bai, Yuntao and
Hatfield-Dodds, Zac and
Clark, Jack and
Bowman, Samuel R. and
Askell, Amanda and
Grosse, Roger and
Hernandez, Danny and
Ganguli, Deep and
Hubinger, Evan and
Schiefer, Nicholas and
Kaplan, Jared",
editor = "Rogers, Anna and
Boyd-Graber, Jordan and
Okazaki, Naoaki",
booktitle = "Findings of the Association for Computational Linguistics: ACL 2023",
month = jul,
year = "2023",
address = "Toronto, Canada",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2023.findings-acl.847/",
doi = "10.18653/v1/2023.findings-acl.847",
pages = "13387--13434"
}