paper-with-me

홈 › Papers

Learning Zero-Sum Linear Quadratic Games with Improved Sample Complexity and Last-Iterate Convergence

2023-09-08 · Jiduan Wu, Anas Barakat, Ilyas Fatkhullin, Niao He

Zero-sum Linear Quadratic (LQ) games are fundamental in optimal control and can be used (i)~as a dynamic game formulation for risk-sensitive or robust control and (ii)~as a benchmark setting for multi-agent reinforcement learning with two competing agents in continuous state-control spaces. In contrast to the well-studied single-agent linear quadratic regulator problem, zero-sum LQ games entail solving a challenging nonconvex-nonconcave min-max problem with an objective function that lacks coercivity. Recently, Zhang et al. showed that an~$\epsilon$-Nash equilibrium (NE) of finite horizon zero-sum LQ games can be learned via nested model-free Natural Policy Gradient (NPG) algorithms with poly$(1/\epsilon)$ sample complexity. In this work, we propose a simpler nested Zeroth-Order (ZO) algorithm improving sample complexity by several orders of magnitude and guaranteeing convergence of the last iterate. Our main results are two-fold: (i) in the deterministic setting, we establish the first global last-iterate linear convergence result for the nested algorithm that seeks NE of zero-sum LQ games; (ii) in the model-free setting, we establish a~$\widetilde{\mathcal{O}}(\epsilon^{-2})$ sample complexity using a single-point ZO estimator. For our last-iterate convergence results, our analysis leverages the Implicit Regularization (IR) property and a new gradient domination condition for the primal function. Our key improvements in the sample complexity rely on a more sample-efficient nested algorithm design and a finer control of the ZO natural gradient estimation error utilizing the structure endowed by the finite-horizon setting.

📄 PDF Abstract BibTeX arXiv:2309.04272

Code (1)

wujiduan/zero-sum-lq-games 공식 구현 pytorch

Tasks

Multi-agent Reinforcement LearningPolicy Gradient Methods

Similar Papers 제목 키워드 기반

Linear-Quadratic Zero-Sum Mean-Field Type Games: Optimality Conditions and Policy Optimization

2020-09-01 · René Carmona, Kenza Hamidouche, Mathieu Laurière, Zongjun Tan

In this paper, zero-sum mean-field type games (ZSMFTG) with linear dynamics and quadratic cost are studied under infinite-horizon discounted utility function. ZSMFTG are a class of games in which two decision makers whos…

Policy Optimization for Linear-Quadratic Zero-Sum Mean-Field Type Games

2020-09-02 · René Carmona, Kenza Hamidouche, Mathieu Laurière, Zongjun Tan

In this paper, zero-sum mean-field type games (ZSMFTG) with linear dynamics and quadratic utility are studied under infinite-horizon discounted utility function. ZSMFTG are a class of games in which two decision makers w…

Vocal Bursts Type Prediction

Covariance steering in zero-sum linear-quadratic two-player differential games

2019-09-12

We formulate a new class of two-person zero-sum differential games, in a stochastic setting, where a specification on a target terminal state distribution is imposed on the players. We address such added specification by…

Vocal Bursts Valence Prediction

Policy Optimization Provably Converges to Nash Equilibria in Zero-Sum Linear Quadratic Games

2019-05-31 · NeurIPS 2019 12 · Kaiqing Zhang, Zhuoran Yang, Tamer Başar

We study the global convergence of policy optimization for finding the Nash equilibria (NE) in zero-sum linear quadratic (LQ) games. To this end, we first investigate the landscape of LQ games, viewing it as a nonconvex-…

Reinforcement Learning

Global Convergence of Policy Gradient for Sequential Zero-Sum Linear Quadratic Dynamic Games

2019-11-12

We propose projection-free sequential algorithms for linear-quadratic dynamics games. These policy gradient based algorithms are akin to Stackelberg leadership model and can be extended to model-free settings. We show th…