Learning Approximate Nash Equilibria in Cooperative Multi-Agent Reinforcement Learning via Mean-Field Subsampling
Many large-scale platforms and networked control systems have a centralized decision maker interacting with a massive population of agents under strict observability constraints. Motivated by such applications, we study a cooperative Markov game with a global agent and $n$ homogeneous local agents in a communication-constrained regime, where the global agent only observes a subset of $k$ local agent states per time step. We propose an alternating learning framework $(\texttt{ALTERNATING-MARL})$, where the global agent performs subsampled mean-field $Q$-learning against a fixed local policy, and local agents update by optimizing in an induced MDP. We prove that these approximate best-response dynamics converge to an $\widetilde{O}(1/\sqrt{k})$-approximate Nash Equilibrium, while separating the sample complexities between the joint state and action spaces. Finally, we validate our results in numerical simulations for multi-robot control.
Code (0)
등록된 구현이 없습니다.
Tasks
Multi-agent Reinforcement LearningSimilar Papers 제목 키워드 기반
NePPO: Near-Potential Policy Optimization for General-Sum Multi-Agent Reinforcement Learning
Multi-agent reinforcement learning (MARL) is increasingly used to design learning-enabled agents that interact in shared environments. However, training MARL algorithms in general-sum games remains challenging: learning …
Multi-agent Reinforcement LearningReinforcement Learning for Finite Space Mean-Field Type Games
Mean field type games (MFTGs) describe Nash equilibria between large coalitions: each coalition consists of a continuum of cooperative agents who maximize the average reward of their coalition while interacting non-coope…
Deep Reinforcement LearningQ-LearningQuantizationreinforcement-learning+1Approximate Nash Equilibrium Learning for n-Player Markov Games in Dynamic Pricing
We investigate Nash equilibrium learning in a competitive Markov Game (MG) environment, where multiple agents compete, and multiple Nash equilibria can exist. In particular, for an oligopolistic dynamic pricing environme…
Q-LearningSolving Nash Equilibria in Nonlinear Differential Games for Common-Pool Resources
Many resources are provided by an ecological system that is vulnerable to tipping when exceeding a certain level of pollution, with a sudden big loss of ecosystem services. An ecological system is usually also a common-p…
MF-OML: Online Mean-Field Reinforcement Learning with Occupation Measures for Large Population Games
Reinforcement learning for multi-agent games has attracted lots of attention recently. However, given the challenge of solving Nash equilibria for large population games, existing works with guaranteed polynomial complex…
Multi-agent Reinforcement Learningreinforcement-learningReinforcement Learning