paper-with-me

홈 › Papers

IPPO Learns the Game, Not the Team: A Study on Generalization in Heterogeneous Agent Teams

2025-12-09 · Ryan LeRoy, Jack Kolb arxiv

Multi-Agent Reinforcement Learning (MARL) is commonly deployed in settings where agents are trained via self-play with homogeneous teammates, often using parameter sharing and a single policy architecture. This opens the question: to what extent do self-play PPO agents learn general coordination strategies grounded in the underlying game, compared to overfitting to their training partners' behaviors? This paper investigates the question using the Heterogeneous Multi-Agent Challenge (HeMAC) environment, which features distinct Observer and Drone agents with complementary capabilities. We introduce Rotating Policy Training (RPT), an approach that rotates heterogeneous teammate policies of different learning algorithms during training, to expose the agent to a broader range of partner strategies. When playing alongside a withheld teammate policy (DDQN), we find that RPT achieves similar performance to a standard self-play baseline, IPPO, where all agents were trained sharing a single PPO policy. This result indicates that in this heterogeneous multi-agent setting, the IPPO baseline generalizes to novel teammate algorithms despite not experiencing teammate diversity during training. This shows that a simple IPPO baseline may possess the level of generalization to novel teammates that a diverse training regimen was designed to achieve.

📄 PDF Abstract BibTeX arXiv:2512.08877

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-agent Reinforcement Learning

Similar Papers 제목 키워드 기반

Transformer Guided Coevolution: Improved Team Selection in Multiagent Adversarial Team Games

2024-10-17 · Pranav Rajbhandari, Prithviraj Dasgupta, Donald Sofge

We consider the problem of team selection within multiagent adversarial team games. We propose BERTeam, a novel algorithm that uses a transformer-based deep neural network with Masked Language Model training to select th…

Deep Reinforcement LearningLanguage ModelingLanguage Modelling

Scheduling Bipartite Tournaments to Minimize Total Travel Distance

2014-01-16 · Richard Hoshino, Ken-ichi Kawarabayashi

In many professional sports leagues, teams from opposing leagues/conferences compete against one another, playing inter-league games. This is an example of a bipartite tournament. In this paper, we consider the problem o…

Scheduling

Super-additive Cooperation in Language Model Agents

2025-08-21 · Filippo Tonini, Lukas Galke arxiv

With the prospect of autonomous artificial intelligence (AI) agents, studying their tendency for cooperative behavior becomes an increasingly relevant topic. This study is inspired by the super-additive cooperation theor…

Role of Externally Provided Randomness in Stochastic Teams and Zero-sum Team Games

2021-10-12 · Rahul Meshram

Stochastic team decision problem is extensively studied in literature and the existence of optimal solution is obtained in recent literature. The value of information in statistical problem and decision theory is classic…

Associative Embedding for Game-Agnostic Team Discrimination

2019-07-01 · Maxime Istasse, Julien Moreau, Christophe De Vleeschouwer

Assigning team labels to players in a sport game is not a trivial task when no prior is known about the visual appearance of each team. Our work builds on a Convolutional Neural Network (CNN) to learn a descriptor, namel…