Value-based CTDE Methods in Symmetric Two-team Markov Game: from Cooperation to Team Competition
In this paper, we identify the best learning scenario to train a team of agents to compete against multiple possible strategies of opposing teams. We evaluate cooperative value-based methods in a mixed cooperative-competitive environment. We restrict ourselves to the case of a symmetric, partially observable, two-team Markov game. We selected three training methods based on the centralised training and decentralised execution (CTDE) paradigm: QMIX, MAVEN and QVMix. For each method, we considered three learning scenarios differentiated by the variety of team policies encountered during training. For our experiments, we modified the StarCraft Multi-Agent Challenge environment to create competitive environments where both teams could learn and compete simultaneously. Our results suggest that training against multiple evolving strategies achieves the best results when, for scoring their performances, teams are faced with several strategies.
Code (2)
Tasks
StarcraftSimilar Papers 제목 키워드 기반
CTDS: Centralized Teacher with Decentralized Student for Multi-Agent Reinforcement Learning
Due to the partial observability and communication constraints in many multi-agent reinforcement learning (MARL) tasks, centralized training with decentralized execution (CTDE) has become one of the most widely used MARL…
Multi-agent Reinforcement Learningreinforcement-learningReinforcement Learning (RL)Starcraft+1Inducing Cooperation via Team Regret Minimization based Multi-Agent Deep Reinforcement Learning
Existing value-factorized based Multi-Agent deep Reinforce-ment Learning (MARL) approaches are well-performing invarious multi-agent cooperative environment under thecen-tralized training and decentralized execution(CTDE…
Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Multi-Agent Deep Reinforcement Learning Under Constrained Communications
Centralized training with decentralized execution (CTDE) has been the dominant paradigm in multi-agent reinforcement learning (MARL), but its reliance on global state information during training introduces scalability, r…
Multi-agent Reinforcement LearningCentralizing State-Values in Dueling Networks for Multi-Robot Reinforcement Learning Mapless Navigation
We study the problem of multi-robot mapless navigation in the popular Centralized Training and Decentralized Execution (CTDE) paradigm. This problem is challenging when each robot considers its path without explicitly sh…
Deep Reinforcement Learningreinforcement-learningReinforcement Learning (RL)RACA: Relation-Aware Credit Assignment for Ad-Hoc Cooperation in Multi-Agent Deep Reinforcement Learning
In recent years, reinforcement learning has faced several challenges in the multi-agent domain, such as the credit assignment issue. Value function factorization emerges as a promising way to handle the credit assignment…
Deep Reinforcement LearningReinforcement Learning (RL)RelationZero-shot Generalization