ICML 2024poster16 citations

STEER: Assessing the Economic Rationality of Large Language Models

Narun Krishnamurthi Raman, Taylor Lundy, Samuel Joseph Amouyal, Yoav Levine, Kevin Leyton-Brown, Moshe Tennenholtz

Abstract

There is increasing interest in using LLMs as decision-making "agents". Doing so includes many degrees of freedom: which model should be used; how should it be prompted; should it be asked to introspect, conduct chain-of-thought reasoning, etc? Settling these questions---and more broadly, determining whether an LLM agent is reliable enough to be trusted---requires a methodology for assessing such an agent's economic rationality. In this paper, we provide one. We begin by surveying the economic literature on rational decision making, taxonomizing a large set of fine-grained "elements" that an agent should exhibit, along with dependencies between them. We then propose a benchmark distribution that quantitatively scores an LLMs performance on these elements and, combined with a user-provided rubric, produces a "rationality report card". Finally, we describe the results of a large-scale empirical experiment with 14 different LLMs, characterizing the both current state of the art and the impact of different model sizes on models' ability to exhibit rational behavior.

BibTeX
@inproceedings{
raman2024steer,
title={{STEER}: Assessing the Economic Rationality of Large Language Models},
author={Narun Krishnamurthi Raman and Taylor Lundy and Samuel Joseph Amouyal and Yoav Levine and Kevin Leyton-Brown and Moshe Tennenholtz},
booktitle={Forty-first International Conference on Machine Learning},
year={2024},
url={https://openreview.net/forum?id=nU1mtFDtMX}
}
STEER: Assessing the Economic Rationality of Large Language Models · ICML 2024