paper-with-me

Papers

Rationality Check! Benchmarking the Rationality of Large Language Models

2025-09-18 · Zhilun Zhou, Jing Yi Wang, Nicholas Sukiennik, Chen Gao, Fengli Xu, Yong Li, James Evans arxiv

Large language models (LLMs), a recent advance in deep learning and machine intelligence, have manifested astonishing capacities, now considered among the most promising for artificial general intelligence. With human-like capabilities, LLMs have been used to simulate humans and serve as AI assistants across many applications. As a result, great concern has arisen about whether and under what circumstances LLMs think and behave like real human agents. Rationality is among the most important concepts in assessing human behavior, both in thinking (i.e., theoretical rationality) and in taking action (i.e., practical rationality). In this work, we propose the first benchmark for evaluating the omnibus rationality of LLMs, covering a wide range of domains and LLMs. The benchmark includes an easy-to-use toolkit, extensive experimental results, and analysis that illuminates where LLMs converge and diverge from idealized human rationality. We believe the benchmark can serve as a foundational tool for both developers and users of LLMs.

📄 PDF Abstract BibTeX arXiv:2509.14546

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Benchmarking the rationality of AI decision making using the transitivity axiom

2025-02-14 · Kiwon Song, James M. Jennings III, Clintin P. Davis-Stober

Fundamental choice axioms, such as transitivity of preference, provide testable conditions for determining whether human decision making is rational, i.e., consistent with a utility representation. Recent work has demons…

BenchmarkingDecision MakingModel SelectionRecommendation Systems

Measuring Stochastic Rationality

2023-03-14 · Efe A. Ok, Gerelt Tserenjigmid

Our goal is to develop a partial ordering method for comparing stochastic choice functions on the basis of their individual rationality. To this end, we assign to any stochastic choice function a one-parameter class of d…

Towards Rationality in Language and Multimodal Agents: A Survey

2024-06-01 · Bowen Jiang, Yangxinyu Xie, Xiaomeng Wang, Yuan Yuan 외

This work discusses how to build more rational language and multimodal agents and what criteria define rationality in intelligent systems. Rationality is the quality of being guided by reason, characterized by decision-m…

Decision MakingSurvey

The Emergence of Economic Rationality of GPT

2023-05-22 · Yiting Chen, Tracy Xiao Liu, You Shan, Songfa Zhong

As large language models (LLMs) like GPT become increasingly prevalent, it is essential that we assess their capabilities beyond language processing. This paper examines the economic rationality of GPT by instructing it …

Governance Challenges in Reinforcement Learning from Human Feedback: Evaluator Rationality and Reinforcement Stability

2025-04-17 · Dana Alsagheer, Abdulrahman Kamal, Mohammad Kamal, Weidong Shi

Reinforcement Learning from Human Feedback (RLHF) is central in aligning large language models (LLMs) with human values and expectations. However, the process remains susceptible to governance challenges, including evalu…

Fairness