A Reputation for Honesty
We analyze situations in which players build reputations for honesty rather than for playing particular actions. A patient player facing a sequence of short-run opponents makes an announcement about their intended action after observing an idiosyncratic shock, and before players act. The patient player is either an honest type whose action coincides with their announcement, or an opportunistic type who can freely choose their actions. We show that the patient player can secure a high payoff by building a reputation for being honest when the short-run players face uncertainty about which of the patient player's actions are currently feasible, but may receive a low payoff when there is no such uncertainty.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
DRF: LLM-AGENT Dynamic Reputation Filtering Framework
With the evolution of generative AI, multi - agent systems leveraging large - language models(LLMs) have emerged as a powerful tool for complex tasks. However, these systems face challenges in quantifying agent performan…
Logical ReasoningPaying for Honesty Without Knowing the Truth: Reputation-Penalty Design for LLM Marketplace Agents
LLM agents increasingly act as autonomous merchants that write their own product listings, and under competitive pressure, they fabricate attributes to win sales. Even under instructions to be honest, they fabricate attr…
Crisis-Bench: Benchmarking Strategic Ambiguity and Reputation Management in Large Language Models
Standard safety alignment optimizes Large Language Models (LLMs) for universal helpfulness and honesty, effectively instilling a rigid "Boy Scout" morality. While robust for general-purpose assistants, this one-size-fits…
Public RelationsAlignment for Honesty
Recent research has made significant strides in aligning large language models (LLMs) with helpfulness and harmlessness. In this paper, we argue for the importance of alignment for \emph{honesty}, ensuring that LLMs proa…
Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning
Unlearning in large language models (LLMs) aims to remove harmful training data while preserving overall utility. However, we find that existing methods often hallucinate, generate abnormal token sequences, or behave inc…