paper-with-me

홈 › Papers

Agent Environment Cycle Games

2020-09-28 · Justin K. Terry, Nathaniel Grammel, Benjamin Black, Ananth Hari, Caroline Horsch, Luis Santos

Partially Observable Stochastic Games (POSGs) are the most general and common model of games used in Multi-Agent Reinforcement Learning (MARL). We argue that the POSG model is conceptually ill suited to software MARL environments, and offer case studies from the literature where this mismatch has led to severely unexpected behavior. In response to this, we introduce the Agent Environment Cycle Games (AEC Games) model, which is more representative of software implementation. We then prove it's as an equivalent model to POSGs. The AEC games model is also uniquely useful in that it can elegantly represent both all forms of MARL environments, whereas for example POSGs cannot elegantly represent strictly turn based games like chess.

📄 PDF Abstract BibTeX arXiv:2009.13051

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-agent Reinforcement Learningreinforcement-learningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

PettingZoo: Gym for Multi-Agent Reinforcement Learning

2020-09-30 · NeurIPS 2021 12 · J. K. Terry, Benjamin Black, Nathaniel Grammel, Mario Jayakumar 외

This paper introduces the PettingZoo library and the accompanying Agent Environment Cycle ("AEC") games model. PettingZoo is a library of diverse sets of multi-agent environments with a universal, elegant Python API. Pet…

Multi-agent Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs

2026-05-28 · Kevin Wang, Anna Thöni, Benjamin Kempinski, Bobby Cheng 외 arxiv

Large language models (LLMs) are increasingly deployed as interactive agents, yet their capacity for social and strategic reasoning over extended interaction remains poorly understood. Existing evaluations rely on static…

Toward Agents That Reason About Their Computation

2025-10-26 · Adrian Orenstein, Jessica Chen, Gwyneth Anne Delos Santos, Bayley Sapara 외 arxiv

While reinforcement learning agents can achieve superhuman performance in many complex tasks, they typically do not become more computationally efficient as they improve. In contrast, humans gradually require less cognit…

Reinforcement Learning

Bandit Learning in Concave N-Person Games

2018-12-01 · NeurIPS 2018 12 · Mario Bravo, David Leslie, Panayotis Mertikopoulos

This paper examines the long-run behavior of learning with bandit feedback in non-cooperative concave games. The bandit framework accounts for extremely low-information environments where the agents may not even know the…

Stochastic Optimization

Bandit learning in concave $N$-person games

2018-10-03 · Mario Bravo, David S. Leslie, Panayotis Mertikopoulos

This paper examines the long-run behavior of learning with bandit feedback in non-cooperative concave games. The bandit framework accounts for extremely low-information environments where the agents may not even know the…

Stochastic Optimization