paper-with-me

Papers

Policy Gradient Methods Find the Nash Equilibrium in N-player General-sum Linear-quadratic Games

2021-07-27 · Ben Hambly, Renyuan Xu, Huining Yang

We consider a general-sum N-player linear-quadratic game with stochastic dynamics over a finite horizon and prove the global convergence of the natural policy gradient method to the Nash equilibrium. In order to prove the convergence of the method, we require a certain amount of noise in the system. We give a condition, essentially a lower bound on the covariance of the noise in terms of the model parameters, in order to guarantee convergence. We illustrate our results with numerical experiments to show that even in situations where the policy gradient method may not converge in the deterministic setting, the addition of noise leads to convergence.

📄 PDF Abstract BibTeX arXiv:2107.13090

Code (0)

등록된 구현이 없습니다.

Tasks

Policy Gradient Methods

Similar Papers 제목 키워드 기반

Approximate Nash Equilibrium Learning for n-Player Markov Games in Dynamic Pricing

2022-07-13 · Larkin Liu

We investigate Nash equilibrium learning in a competitive Markov Game (MG) environment, where multiple agents compete, and multiple Nash equilibria can exist. In particular, for an oligopolistic dynamic pricing environme…

Q-Learning

Equilibrium Selection in Multi-Agent Policy Gradients via Opponent-Aware Basin Entry

2026-05-18 · Yevhen Shcherbinin, Arina Redina, Maxim Kalpin, Vlad Kochetov arxiv

Multi-agent policy-gradient methods have been shown to converge locally near stable Nash equilibria. Local convergence, however, does not determine which equilibrium is reached. We study this question through basin-entry…

Independent Policy Gradient for Large-Scale Markov Potential Games: Sharper Rates, Function Approximation, and Game-Agnostic Convergence

2022-02-08 · Dongsheng Ding, Chen-Yu Wei, Kaiqing Zhang, Mihailo R. Jovanović

We examine global non-asymptotic convergence properties of policy gradient methods for multi-agent reinforcement learning (RL) problems in Markov potential games (MPG). To learn a Nash equilibrium of an MPG in which the …

Multi-agent Reinforcement LearningPolicy Gradient MethodsReinforcement Learning (RL)

Accelerating Nash Learning from Human Feedback via Mirror Prox

2025-05-26 · Daniil Tiapkin, Daniele Calandriello, Denis Belomestny, Eric Moulines 외

Traditional Reinforcement Learning from Human Feedback (RLHF) often relies on reward models, frequently assuming preference structures like the Bradley-Terry model, which may not accurately capture the complexities of re…

Provable Policy Gradient Methods for Average-Reward Markov Potential Games

2024-03-09 · Min Cheng, Ruida Zhou, P. R. Kumar, Chao Tian

We study Markov potential games under the infinite horizon average reward criterion. Most previous studies have been for discounted rewards. We prove that both algorithms based on independent policy gradient and independ…

Policy Gradient Methods