EMNLP 2024main0 citations

The Greatest Good Benchmark: Measuring LLMs’ Alignment with Utilitarian Moral Dilemmas

Giovanni Franco Gabriel Marraffini, Andrés Cotton, Noe Fabian Hsueh, Axel Fridman, Juan Wisznia, Luciano Del Corro

Abstract

The question of how to make decisions that maximise the well-being of all persons is very relevant to design language models that are beneficial to humanity and free from harm. We introduce the Greatest Good Benchmark to evaluate the moral judgments of LLMs using utilitarian dilemmas. Our analysis across 15 diverse LLMs reveals consistently encoded moral preferences that diverge from established moral theories and lay population moral standards. Most LLMs have a marked preference for impartial beneficence and rejection of instrumental harm. These findings showcase the ‘artificial moral compass’ of LLMs, offering insights into their moral alignment.

BibTeX
@inproceedings{marraffini-etal-2024-greatest,
    title = "The Greatest Good Benchmark: Measuring {LLM}s' Alignment with Utilitarian Moral Dilemmas",
    author = "Marraffini, Giovanni Franco Gabriel  and
      Cotton, Andr{\'e}s  and
      Hsueh, Noe Fabian  and
      Fridman, Axel  and
      Wisznia, Juan  and
      Corro, Luciano Del",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.1224/",
    doi = "10.18653/v1/2024.emnlp-main.1224",
    pages = "21950--21959"
}