paper-with-me

홈 › Papers

Welfare Diplomacy: Benchmarking Language Model Cooperation

2023-10-13 · Gabriel Mukobi, Hannah Erlebach, Niklas Lauffer, Lewis Hammond, Alan Chan, Jesse Clifton

The growing capabilities and increasingly widespread deployment of AI systems necessitate robust benchmarks for measuring their cooperative capabilities. Unfortunately, most multi-agent benchmarks are either zero-sum or purely cooperative, providing limited opportunities for such measurements. We introduce a general-sum variant of the zero-sum board game Diplomacy -- called Welfare Diplomacy -- in which players must balance investing in military conquest and domestic welfare. We argue that Welfare Diplomacy facilitates both a clearer assessment of and stronger training incentives for cooperative capabilities. Our contributions are: (1) proposing the Welfare Diplomacy rules and implementing them via an open-source Diplomacy engine; (2) constructing baseline agents using zero-shot prompted language models; and (3) conducting experiments where we find that baselines using state-of-the-art models attain high social welfare but are exploitable. Our work aims to promote societal safety by aiding researchers in developing and assessing multi-agent AI systems. Code to evaluate Welfare Diplomacy and reproduce our experiments is available at https://github.com/mukobi/welfare-diplomacy.

📄 PDF Abstract BibTeX arXiv:2310.08901

Code (1)

mukobi/welfare-diplomacy 공식 구현 pytorch

Tasks

BenchmarkingLanguage ModelingLanguage Modellingmodel

Similar Papers 제목 키워드 기반

Human-Level Performance in No-Press Diplomacy via Equilibrium Search

2020-10-06 · ICLR 2021 1 · Jonathan Gray, Adam Lerer, Anton Bakhtin, Noam Brown

Prior AI breakthroughs in complex games have focused on either the purely adversarial or purely cooperative settings. In contrast, Diplomacy is a game of shifting alliances that involves both cooperation and competition.…

Social welfare optimisation in well-mixed and structured populations

2025-12-08 · Van An Nguyen, Vuong Khang Huynh, Ho Nam Duong, Huu Loi Bui 외 arxiv

Research on promoting cooperation among autonomous, self-regarding agents has often focused on the bi-objective optimisation problem: minimising the total incentive cost while maximising the frequency of cooperation. How…

Human-level play in the game of Diplomacy by combining language models with strategic reasoning

2022-11-22 · Science 2022 11 · Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina 외

Despite much progress in training AI systems to imitate human language, building agents that use language to communicate intentionally with humans in interactive environments remains a major challenge. We introduce Cicer…

AI AgentLanguage Modeling

More Victories, Less Cooperation: Assessing Cicero's Diplomacy Play

2024-06-07 · Wichayaporn Wongkamjan, Feng Gu, Yanze Wang, Ulf Hermjakob 외

The boardgame Diplomacy is a challenging setting for communicative and cooperative artificial intelligence. The most prominent communicative Diplomacy AI, Cicero, has excellent strategic abilities, exceeding human player…

Abstract Meaning Representation

Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning

2022-10-11 · Anton Bakhtin, David J Wu, Adam Lerer, Jonathan Gray 외

No-press Diplomacy is a complex strategy game involving both cooperation and competition that has served as a benchmark for multi-agent AI research. While self-play reinforcement learning has resulted in numerous success…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)