paper-with-me

Papers

STEER: Assessing the Economic Rationality of Large Language Models

2024-02-14 · Narun Raman, Taylor Lundy, Samuel Amouyal, Yoav Levine, Kevin Leyton-Brown, Moshe Tennenholtz

There is increasing interest in using LLMs as decision-making "agents." Doing so includes many degrees of freedom: which model should be used; how should it be prompted; should it be asked to introspect, conduct chain-of-thought reasoning, etc? Settling these questions -- and more broadly, determining whether an LLM agent is reliable enough to be trusted -- requires a methodology for assessing such an agent's economic rationality. In this paper, we provide one. We begin by surveying the economic literature on rational decision making, taxonomizing a large set of fine-grained "elements" that an agent should exhibit, along with dependencies between them. We then propose a benchmark distribution that quantitatively scores an LLMs performance on these elements and, combined with a user-provided rubric, produces a "STEER report card." Finally, we describe the results of a large-scale empirical experiment with 14 different LLMs, characterizing the both current state of the art and the impact of different model sizes on models' ability to exhibit rational behavior.

📄 PDF Abstract BibTeX arXiv:2402.09552

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Making

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

The Emergence of Economic Rationality of GPT

2023-05-22 · Yiting Chen, Tracy Xiao Liu, You Shan, Songfa Zhong

As large language models (LLMs) like GPT become increasingly prevalent, it is essential that we assess their capabilities beyond language processing. This paper examines the economic rationality of GPT by instructing it …

Sparks of Rationality: Do Reasoning LLMs Align with Human Judgment and Choice?

2026-01-29 · Ala N. Tak, Amin Banayeeanzade, Anahita Bolourani, Fatemeh Bahrani 외 arxiv

Large Language Models (LLMs) are increasingly positioned as decision engines for hiring, healthcare, and economic judgment, yet real-world human judgment reflects a balance between rational deliberation and emotion-drive…

STEER-ME: Assessing the Microeconomic Reasoning of Large Language Models

2025-02-18 · Narun Raman, Taylor Lundy, Thiago Amin, Jesse Perla 외

How should one judge whether a given large language model (LLM) can reliably perform economic reasoning? Most existing LLM benchmarks focus on specific applications and fail to present the model with a rich variety of ec…

BenchmarkingLarge Language Model

Economic Rationality under Specialization: Evidence of Decision Bias in AI Agents

2025-01-30 · ShuiDe Wen, Juan Feng

In the study by Chen et al. (2023) [01], the large language model GPT demonstrated economic rationality comparable to or exceeding the average human level in tasks such as budget allocation and risk preference. Building …

Decision MakingLanguage ModelingLanguage ModellingLarge Language Model

Rationality Check! Benchmarking the Rationality of Large Language Models

2025-09-18 · Zhilun Zhou, Jing Yi Wang, Nicholas Sukiennik, Chen Gao 외 arxiv

Large language models (LLMs), a recent advance in deep learning and machine intelligence, have manifested astonishing capacities, now considered among the most promising for artificial general intelligence. With human-li…