Moral Hazard in Multi-Agent Language Models
Cooperation can fail when socially valuable effort is costly, hard to observe, and benefits mainly someone else. Building on Holmström's model of moral hazard in teams, we introduce the Dialogue Moral Hazard Game, a theory-grounded controlled experimental paradigm that instantiates this hidden-action structure as a textual environment for language agents. In each episode, an agent chooses between keeping an immediate local reward and paying a query cost to reveal a hidden safety fact that primarily helps another agent's downstream decision. We evaluate eleven open-weight language models and three frontier API models, decomposing behavior into query rate, realized information transfer, local-reward preservation, unsafe choice, format validity, and team success. The frontier policies differ sharply: Fable 5 moves from querying toward local reward as cost rises and back toward querying as team reward rises, yet remains query-saturated under controlled private-share isolation; Muse Spark 1.1 responds to query cost, team reward, and private team share; and GPT-5.6 Sol reaches ceiling behavior in the primary setting. In a 3,015-decision incentive-isolation experiment, Sol tracks the Holmström-derived private-share boundary across nine query costs with a mean absolute error of 0.013. We then apply supervised fine-tuning, RLOO, sequential SFT+RLOO, and GEPA prompt optimization as diagnostic update mechanisms wherever model access permits. Their effects are heterogeneous: SmolLM3-3B and OLMo-7B show the clearest mechanism-consistent, weight-level gains, whereas GEPA sometimes raises team success while reducing or eliminating costly queries. Optimization can therefore lift aggregate reward without restoring the designated cooperative mechanism, motivating evaluations that report mechanism-level behavior rather than team success alone.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Optimism, overconfidence, and moral hazard
I revisit the standard moral-hazard model, in which an agent's preference over contracts is rooted in costly effort choice. I characterise the behavioural content of the model in terms of empirically testable axioms, and…
Moral Hazard, Dynamic Incentives, and Ambiguous Perceptions
This paper considers dynamic moral hazard settings, in which the consequences of the agent's actions are not precisely understood. In a new continuous-time moral hazard model with drift ambiguity, the agent's unobservabl…
Ex-post moral hazard and manipulation-proof contracts
We examine the trade-off between the provision of incentives to exert costly effort (ex-ante moral hazard) and the incentives needed to prevent the agent from manipulating the profit observed by the principal (ex-post mo…
Moral Hazard in Dynamic Risk Management
We consider a contracting problem in which a principal hires an agent to manage a risky project. When the agent chooses volatility components of the output process and the principal observes the output continuously, the …
ManagementOptimal contracts under competition when uncertainty from adverse selection and moral hazard are present
In a continuous-time setting where a risk-averse agent controls the drift of an output process driven by a Brownian motion, optimal contracts are linear in the terminal output; this result is well-known in a setting with…