STAR: SocioTechnical Approach to Red Teaming Language Models
Laura Weidinger, John F J Mellor, Bernat Guillén Pegueroles, Nahema Marchal, Ravin Kumar, Kristian Lum, Canfer Akbulut, Mark Diaz
Abstract
This research introduces STAR, a sociotechnical framework that improves on current best practices for red teaming safety of large language models. STAR makes two key contributions: it enhances steerability by generating parameterised instructions for human red teamers, leading to improved coverage of the risk surface. Parameterised instructions also provide more detailed insights into model failures at no increased cost. Second, STAR improves signal quality by matching demographics to assess harms for specific groups, resulting in more sensitive annotations. STAR further employs a novel step of arbitration to leverage diverse viewpoints and improve label reliability, treating disagreement not as noise but as a valuable contribution to signal quality.
BibTeX
@inproceedings{weidinger-etal-2024-star,
title = "{STAR}: {S}ocio{T}echnical Approach to Red Teaming Language Models",
author = "Weidinger, Laura and
Mellor, John F J and
Pegueroles, Bernat Guill{\'e}n and
Marchal, Nahema and
Kumar, Ravin and
Lum, Kristian and
Akbulut, Canfer and
Diaz, Mark and
Bergman, A. Stevie and
Rodriguez, Mikel D. and
Rieser, Verena and
Isaac, William",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen, Yun-Nung",
booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
month = nov,
year = "2024",
address = "Miami, Florida, USA",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.emnlp-main.1200/",
doi = "10.18653/v1/2024.emnlp-main.1200",
pages = "21516--21532"
}