paper-with-me

홈 › Papers

Learning to Resolve Alliance Dilemmas in Many-Player Zero-Sum Games

2020-02-27 · Edward Hughes, Thomas W. Anthony, Tom Eccles, Joel Z. Leibo, David Balduzzi, Yoram Bachrach

Zero-sum games have long guided artificial intelligence research, since they possess both a rich strategy space of best-responses and a clear evaluation metric. What's more, competition is a vital mechanism in many real-world multi-agent systems capable of generating intelligent innovations: Darwinian evolution, the market economy and the AlphaZero algorithm, to name a few. In two-player zero-sum games, the challenge is usually viewed as finding Nash equilibrium strategies, safeguarding against exploitation regardless of the opponent. While this captures the intricacies of chess or Go, it avoids the notion of cooperation with co-players, a hallmark of the major transitions leading from unicellular organisms to human civilization. Beyond two players, alliance formation often confers an advantage; however this requires trust, namely the promise of mutual cooperation in the face of incentives to defect. Successful play therefore requires adaptation to co-players rather than the pursuit of non-exploitability. Here we argue that a systematic study of many-player zero-sum games is a crucial element of artificial intelligence research. Using symmetric zero-sum matrix games, we demonstrate formally that alliance formation may be seen as a social dilemma, and empirically that na\"ive multi-agent reinforcement learning therefore fails to form alliances. We introduce a toy model of economic competition, and show how reinforcement learning may be augmented with a peer-to-peer contract mechanism to discover and enforce alliances. Finally, we generalize our agent model to incorporate temporally-extended contracts, presenting opportunities for further work.

📄 PDF Abstract BibTeX arXiv:2003.00799

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-agent Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

AlphaZero AlphaZero is a reinforcement learning agent for playing board games such as Go, chess, and shogi.

Similar Papers 제목 키워드 기반

Learning to Play No-Press Diplomacy with Best Response Policy Iteration

2020-06-08 · NeurIPS 2020 12 · Thomas Anthony, Tom Eccles, Andrea Tacchetti, János Kramár 외

Recent advances in deep reinforcement learning (RL) have led to considerable progress in many 2-player zero-sum games, such as Go, Poker and Starcraft. The purely adversarial nature of such games allows for conceptually …

Deep Reinforcement LearningReinforcement Learning (RL)Starcraft

Learning Reciprocity in Complex Sequential Social Dilemmas

2019-03-19 · Tom Eccles, Edward Hughes, János Kramár, Steven Wheelwright 외

Reciprocity is an important feature of human social interaction and underpins our cooperative nature. What is more, simple forms of reciprocity have proved remarkably resilient in matrix game social dilemmas. Most famous…

Reinforcement Learning

It Takes Two to Lie: One to Lie, and One to Listen

2020-07-01 · ACL 2020 6 · Denis Peskov, Benny Cheng, Ahmed Elgohary, Joe Barrow 외

Trust is implicit in many online text conversations{---}striking up new friendships, or asking for tech support. But trust can be betrayed through deception. We study the language and dynamics of deception in the negotia…

Evolutionary dynamics of zero-determinant strategies in repeated multiplayer games

2021-09-14 · Fang Chen, Te Wu, Long Wang

Since Press and Dyson's ingenious discovery of ZD (zero-determinant) strategy in the repeated Prisoner's Dilemma game, several studies have confirmed the existence of ZD strategy in repeated multiplayer social dilemmas. …

Heterogeneous Social Value Orientation Leads to Meaningful Diversity in Sequential Social Dilemmas

2023-05-01 · Udari Madhushani, Kevin R. McKee, John P. Agapiou, Joel Z. Leibo 외

In social psychology, Social Value Orientation (SVO) describes an individual's propensity to allocate resources between themself and others. In reinforcement learning, SVO has been instantiated as an intrinsic motivation…

DiversityZero-shot Generalization