paper-with-me

홈 › Papers

Asymptotic Convergence and Performance of Multi-Agent Q-Learning Dynamics

2023-01-23 · Aamal Abbas Hussain, Francesco Belardinelli, Georgios Piliouras

Achieving convergence of multiple learning agents in general $N$-player games is imperative for the development of safe and reliable machine learning (ML) algorithms and their application to autonomous systems. Yet it is known that, outside the bounds of simple two-player games, convergence cannot be taken for granted. To make progress in resolving this problem, we study the dynamics of smooth Q-Learning, a popular reinforcement learning algorithm which quantifies the tendency for learning agents to explore their state space or exploit their payoffs. We show a sufficient condition on the rate of exploration such that the Q-Learning dynamics is guaranteed to converge to a unique equilibrium in any game. We connect this result to games for which Q-Learning is known to converge with arbitrary exploration rates, including weighted Potential games and weighted zero sum polymatrix games. Finally, we examine the performance of the Q-Learning dynamic as measured by the Time Averaged Social Welfare, and comparing this with the Social Welfare achieved by the equilibrium. We provide a sufficient condition whereby the Q-Learning dynamic will outperform the equilibrium even if the dynamics do not converge.

📄 PDF Abstract BibTeX arXiv:2301.09619

Code (0)

등록된 구현이 없습니다.

Tasks

Q-Learning

Methods 이 논문이 사용한 방법론

Q-Learning Q-Learning is an off-policy temporal difference control algorithm: $$Q\left(S\_{t}, A\_{t}\right) \leftarrow Q\left(S\_{t}, A\_{t}\right) + \alpha\left[R_{t+1} +…

Similar Papers 제목 키워드 기반

Does Worst-Performing Agent Lead the Pack? Analyzing Agent Dynamics in Unified Distributed SGD

2024-09-26 · Jie Hu, Yi-Ting Ma, Do Young Eun

Distributed learning is essential to train machine learning algorithms across heterogeneous agents while maintaining data privacy. We conduct an asymptotic analysis of Unified Distributed SGD (UD-SGD), exploring a variet…

Federated Learning

Convergence Analysis of Weighted-Median Opinion Dynamics with Prejudice

2024-08-04 · Ruichang Zhang, Zhixin Liu, Ge Chen, Wenjun Mei

The Friedkin-Johnsen (FJ) model introduces prejudice into the opinion evolution and has been successfully validated in many practical scenarios; however, due to its weighted average mechanism, only one prejudiced agent c…

Distributed Mirror Descent with Integral Feedback: Asymptotic Convergence Analysis of Continuous-time Dynamics

2020-09-14 · Youbang Sun, Shahin Shahrampour

This work addresses distributed optimization, where a network of agents wants to minimize a global strongly convex objective function. The global function can be written as a sum of local convex functions, each of which …

Distributed Optimization

Model-based Multi-agent Policy Optimization with Adaptive Opponent-wise Rollouts

2021-05-07 · Weinan Zhang, Xihuai Wang, Jian Shen, Ming Zhou

This paper investigates the model-based methods in multi-agent reinforcement learning (MARL). We specify the dynamics sample complexity and the opponent sample complexity in MARL, and conduct a theoretic analysis of retu…

Multi-agent Reinforcement LearningReinforcement Learning (RL)

Friedkin-Johnsen Model with Diminishing Competition

2024-09-19 · Luca Ballotta, Áron Vékássy, Stephanie Gil, Michal Yemini

This letter studies the Friedkin-Johnsen (FJ) model with diminishing competition, or stubbornness. The original FJ model assumes that each agent assigns a constant competition weight to its initial opinion. In contrast, …

model