paper-with-me

Papers

Multi-Agent Reinforcement Learning in Cournot Games

2020-09-14 · Yuanyuan Shi, Baosen Zhang

In this work, we study the interaction of strategic agents in continuous action Cournot games with limited information feedback. Cournot game is the essential market model for many socio-economic systems where agents learn and compete without the full knowledge of the system or each other. We consider the dynamics of the policy gradient algorithm, which is a widely adopted continuous control reinforcement learning algorithm, in concave Cournot games. We prove the convergence of policy gradient dynamics to the Nash equilibrium when the price function is linear or the number of agents is two. This is the first result (to the best of our knowledge) on the convergence property of learning algorithms with continuous action spaces that do not fall in the no-regret class.

📄 PDF Abstract BibTeX arXiv:2009.06224

Code (0)

등록된 구현이 없습니다.

Tasks

continuous-controlContinuous ControlMulti-agent Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Using Non-Stationary Bandits for Learning in Repeated Cournot Games with Non-Stationary Demand

2022-01-03 · Kshitija Taywade, Brent Harrison, Judy Goldsmith

Many past attempts at modeling repeated Cournot games assume that demand is stationary. This does not align with real-world scenarios in which market demands can evolve over a product's lifetime for a myriad of reasons. …

Efficient Exploration

Modelling Cournot Games as Multi-agent Multi-armed Bandits

2022-01-01 · Kshitija Taywade, Brent Harrison, Adib Bagh

We investigate the use of a multi-agent multi-armed bandit (MA-MAB) setting for modeling repeated Cournot oligopoly games, where the firms acting as agents choose from the set of arms representing production quantity (a …

Multi-Armed Bandits

On the Oscillations in Cournot Games with Best Response Strategies

2024-10-12 · Zhengyang Liu, Haolin Lu, Liang Shan, Zihe Wang

In this paper, we consider the dynamic oscillation in the Cournot oligopoly model, which involves multiple firms producing homogeneous products. To explore the oscillation under the updates of best response strategies, w…

Gradient-Tracking over Directed Graphs for solving Leaderless Multi-Cluster Games

2021-02-18 · Jan Zimmermann, Tatiana Tatarenko, Volker Willert, Jürgen Adamy

We are concerned with finding Nash Equilibria in agent-based multi-cluster games, where agents are separated into distinct clusters. While the agents inside each cluster collaborate to achieve a common goal, the clusters…

Distributed Optimization

A Zeroth-Order Momentum Method for Risk-Averse Online Convex Games

2022-09-06 · Zifan Wang, Yi Shen, Zachary I. Bell, Scott Nivison 외

We consider risk-averse learning in repeated unknown games where the goal of the agents is to minimize their individual risk of incurring significantly high cost. Specifically, the agents use the conditional value at ris…