paper-with-me

홈 › Papers

Learning Cooperative Multi-Agent Policies with Partial Reward Decoupling

2021-12-23 · Benjamin Freed, Aditya Kapoor, Ian Abraham, Jeff Schneider, Howie Choset

One of the preeminent obstacles to scaling multi-agent reinforcement learning to large numbers of agents is assigning credit to individual agents' actions. In this paper, we address this credit assignment problem with an approach that we call \textit{partial reward decoupling} (PRD), which attempts to decompose large cooperative multi-agent RL problems into decoupled subproblems involving subsets of agents, thereby simplifying credit assignment. We empirically demonstrate that decomposing the RL problem using PRD in an actor-critic algorithm results in lower variance policy gradient estimates, which improves data efficiency, learning stability, and asymptotic performance across a wide array of multi-agent RL tasks, compared to various other actor-critic approaches. Additionally, we relate our approach to counterfactual multi-agent policy gradient (COMA), a state-of-the-art MARL algorithm, and empirically show that our approach outperforms COMA by making better use of information in agents' reward streams, and by enabling recent advances in advantage estimation to be used.

📄 PDF Abstract BibTeX arXiv:2112.12740

Code (0)

등록된 구현이 없습니다.

Tasks

counterfactualMulti-agent Reinforcement Learning

Similar Papers 제목 키워드 기반

Learning Reward Machines in Cooperative Multi-Agent Tasks

2023-03-24 · Leo Ardon, Daniel Furelos-Blanco, Alessandra Russo

This paper presents a novel approach to Multi-Agent Reinforcement Learning (MARL) that combines cooperative task decomposition with the learning of reward machines (RMs) encoding the structure of the sub-tasks. The propo…

Multi-agent Reinforcement Learning

Reward Design in Cooperative Multi-agent Reinforcement Learning for Packet Routing

2020-03-05 · ICLR 2018 1 · Hangyu Mao, Zhibo Gong, Zhen Xiao

In cooperative multi-agent reinforcement learning (MARL), how to design a suitable reward signal to accelerate learning and stabilize convergence is a critical problem. The global reward signal assigns the same global re…

Multi-agent Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Hybrid Differential Reward: Combining Temporal Difference and Action Gradients for Efficient Multi-Agent Reinforcement Learning in Cooperative Driving

2025-11-21 · Ye Han, Lijun Zhang, Dejian Meng, Zhuang Zhang arxiv

In multi-vehicle cooperative driving tasks involving high-frequency continuous control, traditional state-based reward functions suffer from the issue of vanishing reward differences. This phenomenon results in a low sig…

Multi-agent Reinforcement LearningContinuous Control

MACRPO: Multi-Agent Cooperative Recurrent Policy Optimization

2021-09-02 · Eshagh Kargar, Ville Kyrki

This work considers the problem of learning cooperative policies in multi-agent settings with partially observable and non-stationary environments without a communication channel. We focus on improving information sharin…

HAVEN: Hierarchical Cooperative Multi-Agent Reinforcement Learning with Dual Coordination Mechanism

2021-10-14 · Zhiwei Xu, Yunpeng Bai, Bin Zhang, Dapeng Li 외

Recently, some challenging tasks in multi-agent systems have been solved by some hierarchical reinforcement learning methods. Inspired by the intra-level and inter-level coordination in the human nervous system, we propo…

Hierarchical Reinforcement LearningMulti-agent Reinforcement Learningreinforcement-learningReinforcement Learning+3