paper-with-me

홈 › Papers

When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games

2026-07-06 · Jerick Shi, Terry Jingcheng Zhang, Bernhard Schölkopf, Vincent Conitzer, Zhijing Jin arxiv

As large language models are deployed as autonomous agents that communicate intentions before acting, a critical safety question is whether agents that publicly commit to actions will honor those commitments. We place LLM agents in repeated $n$-player games with a three-stage protocol that separates private intent, public announcement, and final action, allowing us to identify whether each deviation from a stated announcement was already planned during private deliberation. Evaluating three frontier models across six games in homogeneous and heterogeneous groups over 10 rounds, we report two findings. First, when agents deviate from their announcements, the deviation is predominantly already stated in their private plan (exceeding 90% in the highest-deception conditions), yet this is not a fixed model property: the same model ranges from perfect honesty to near-total deviation across games. Second, different models interpret announcements incompatibly, some as binding commitments and others as cheap talk, producing payoff gaps that emerge in Round~0 and persist across all 10 rounds. Systems that combine models from different providers therefore cannot assume shared announcement semantics and require empirical testing of model interactions before deployment.

📄 PDF Abstract BibTeX arXiv:2607.05132

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Disentangling Exploration from Exploitation

2024-04-29 · Alessandro Lizzeri, Eran Shmaya, Leeat Yariv

Starting from Robbins (1952), the literature on experimentation via multi-armed bandits has wed exploration and exploitation. Nonetheless, in many applications, agents' exploration and exploitation need not be intertwine…

DisentanglementMulti-Armed Bandits

Push Your Agent: Measuring and Enforcing Quantitative Goal Persistence in Long-Horizon LLM Agents

2026-05-22 · Yuandao Cai, Yuzhang Zhu, Liyou Gao, Wensheng Tang 외 arxiv

Long-horizon language agents can make many plausible local tool calls yet fail to persist until a requested count is actually complete. We study this gap as Quantitative Goal Persistence (QGP): whether an agent keeps wor…

State-Novelty Guided Action Persistence in Deep Reinforcement Learning

2024-09-09 · Jianshu Hu, Paul Weng, Yutong Ban

While a powerful and promising approach, deep reinforcement learning (DRL) still suffers from sample inefficiency, which can be notably improved by resorting to more sophisticated techniques to address the exploration-ex…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningScheduling

Evolutionarily Stable (Mis)specifications: Theory and Applications

2020-12-30 · Kevin He, Jonathan Libgober

Toward explaining the persistence of biased inferences, we propose a framework to evaluate competing (mis)specifications in strategic settings. Agents with heterogeneous (mis)specifications coexist and draw Bayesian infe…

Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows

2026-04-22 · Hardy Chen, Nancy Lau, Haoqin Tu, Shuo Yan 외 arxiv

Frontier coding agents are increasingly used in workflows where users supervise progress primarily through repeated improvement of a public score, namely the reported score on a public evaluation file with labels in the …