paper-with-me

홈 › Papers

Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation

2026-06-16 · Yuyang Dai, Xueqing Peng, Lingfei Qian, Zhuohan Xie arxiv

Evaluating the decision-making capabilities of large language models (LLMs) is a growing research priority, yet existing benchmarks focus on isolated cognitive tasks such as reasoning, knowledge retrieval, and economic rationality in stylized settings. These evaluations overlook the defining challenge of real executive decision-making: integrating conflicting recommendations from specialized stakeholders under information asymmetry, organizational constraints, and temporal dependencies. We introduce \textsc{CEO-Bench}, a multi-agent benchmark that evaluates LLMs on CEO-level strategic resource reallocation -- the process of redirecting capital across business units in a multi-round, constraint-rich organizational environment. In \textsc{CEO-Bench}, LLM agents receive conflicting advice from four role-conditioned C-suite advisors (CFO, CTO, COO, CMO), each with private signals and distinct priorities, and must synthesize these into a concrete allocation plan evaluated along four dimensions: role integration, conditional boldness, history-sensitive judgment, and plan validity. Experiments across five frontier models on 13 scenarios reveal that all models achieve high structural validity but diverge sharply on strategic calibration -- the hardest capability layer. We identify systematic failure modes including single-advisor capture, conservative default under ambiguity, and historical amnesia, and uncover a structural integration-boldness tradeoff: models that engage more deeply with conflicting perspectives tend to produce less decisive action. These findings delineate the current capability boundary of LLMs as organizational decision-makers and inform the design of future AI-assisted executive systems.

📄 PDF Abstract BibTeX arXiv:2606.17459

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Variable population manipulations of reallocation rules in economies with single-peaked preferences

2022-10-23 · Agustin G. Bonifacio

In a one-commodity economy with single-peaked preferences and individual endowments, we study different ways in which reallocation rules can be strategically distorted by affecting the set of active agents. We introduce …

Computation Reallocation for Object Detection

2019-12-24 · ICLR 2020 1 · Feng Liang, Chen Lin, Ronghao Guo, Ming Sun 외

The allocation of computation resources in the backbone is a crucial issue in object detection. However, classification allocation pattern is usually adopted directly to object detector, which is proved to be sub-optimal…

Instance SegmentationNeural Architecture SearchObjectobject-detection+2

Stanceosaurus: Classifying Stance Towards Multilingual Misinformation

2022-10-28 · Jonathan Zheng, Ashutosh Baheti, Tarek Naous, Wei Xu 외

We present Stanceosaurus, a new corpus of 28,033 tweets in English, Hindi, and Arabic annotated with stance towards 251 misinformation claims. As far as we are aware, it is the largest corpus annotated with stance toward…

Domain AdaptationFact CheckingMisinformation

Do You Get the Hint? Benchmarking LLMs on the Board Game Concept

2025-10-15 · Ine Gevers, Walter Daelemans arxiv

Large language models (LLMs) have achieved striking successes on many benchmarks, yet recent studies continue to expose fundamental weaknesses. In this paper, we introduce Concept, a simple word-guessing board game, as a…

Network Pruning via Resource Reallocation

2021-03-02 · Yuenan Hou, Zheng Ma, Chunxiao Liu, Zhe Wang 외

Channel pruning is broadly recognized as an effective approach to obtain a small compact model through eliminating unimportant channels from a large cumbersome network. Contemporary methods typically perform iterative pr…

Network Pruning