paper-with-me

Papers

Tackling Asymmetric and Circular Sequential Social Dilemmas with Reinforcement Learning and Graph-based Tit-for-Tat

2022-06-26 · Tangui Le Gléau, Xavier Marjou, Tayeb Lemlouma, Benoit Radier

In many societal and industrial interactions, participants generally prefer their pure self-interest at the expense of the global welfare. Known as social dilemmas, this category of non-cooperative games offers situations where multiple actors should all cooperate to achieve the best outcome but greed and fear lead to a worst self-interested issue. Recently, the emergence of Deep Reinforcement Learning (RL) has generated revived interest in social dilemmas with the introduction of Sequential Social Dilemma (SSD). Cooperative agents mixing RL policies and Tit-for-tat (TFT) strategies have successfully addressed some non-optimal Nash equilibrium issues. However, this kind of paradigm requires symmetrical and direct cooperation between actors, conditions that are not met when mutual cooperation become asymmetric and is possible only with at least a third actor in a circular way. To tackle this issue, this paper extends SSD with Circular Sequential Social Dilemma (CSSD), a new kind of Markov games that better generalizes the diversity of cooperation between agents. Secondly, to address such circular and asymmetric cooperation, we propose a candidate solution based on RL policies and a graph-based TFT. We conducted some experiments on a simple multi-player grid world which offers adaptable cooperation structures. Our work confirmed that our graph-based approach is beneficial to address circular situations by encouraging self-interested agents to reach mutual cooperation.

📄 PDF Abstract BibTeX arXiv:2206.12909

Code (1)

submission-conf/neurips_cooperativeai 공식 구현 pytorch

Tasks

Deep Reinforcement LearningReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Non Maximum Suppression Non Maximum Suppression is a computer vision method that selects a single entity out of many overlapping entities (for example bounding boxes in object detection). The…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
SSD SSD is a single-stage object detection method that discretizes the output space of bounding boxes into a set of default boxes over different aspect ratios and scales per…

Similar Papers 제목 키워드 기반

Fairness over Equality: Correcting Social Incentives in Asymmetric Sequential Social Dilemmas

2026-02-17 · Alper Demir, Hüseyin Aydın, Kale-ab Abebe Tessera, David Abel 외 arxiv

Sequential Social Dilemmas (SSDs) provide a key framework for studying how cooperation emerges when individual incentives conflict with collective welfare. In Multi-Agent Reinforcement Learning, these problems are often …

Multi-agent Reinforcement Learning

Multi-agent Reinforcement Learning in Sequential Social Dilemmas

2017-02-10 · Joel Z. Leibo, Vinicius Zambaldi, Marc Lanctot, Janusz Marecki 외

Matrix games like Prisoner's Dilemma have guided research on social dilemmas for decades. However, they necessarily treat the choice to cooperate or defect as an atomic action. In real-world social dilemmas these choices…

Multi-agent Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Inducing Cooperative behaviour in Sequential-Social dilemmas through Multi-Agent Reinforcement Learning using Status-Quo Loss

2020-01-15 · Pinkesh Badjatiya, Mausoom Sarkar, Abhishek Sinha, Siddharth Singh 외

In social dilemma situations, individual rationality leads to sub-optimal group outcomes. Several human engagements can be modeled as a sequential (multi-step) social dilemmas. However, in contrast to humans, Deep Reinfo…

ClusteringDeep Reinforcement LearningMulti-agent Reinforcement LearningReinforcement Learning

Learning Homophilic Incentives in Sequential Social Dilemmas

2021-09-29 · Heng Dong, Tonghan Wang, Jiayuan Liu, Chi Han 외

Promoting cooperation among self-interested agents is a long-standing and interdisciplinary problem, but receives less attention in multi-agent reinforcement learning (MARL). Game-theoretical studies reveal that altruist…

Multi-agent Reinforcement Learning

SocialJax: An Evaluation Suite for Multi-agent Reinforcement Learning in Sequential Social Dilemmas

2025-03-18 · Zihao Guo, Shuqing Shi, Richard Willis, Tristan Tomilin 외

Sequential social dilemmas pose a significant challenge in the field of multi-agent reinforcement learning (MARL), requiring environments that accurately reflect the tension between individual and collective interests. P…

Multi-agent Reinforcement Learningreinforcement-learningReinforcement Learningrllib