paper-with-me

홈 › Papers

Policy Optimization for Linear-Quadratic Zero-Sum Mean-Field Type Games

2020-09-02 · René Carmona, Kenza Hamidouche, Mathieu Laurière, Zongjun Tan

In this paper, zero-sum mean-field type games (ZSMFTG) with linear dynamics and quadratic utility are studied under infinite-horizon discounted utility function. ZSMFTG are a class of games in which two decision makers whose utilities sum to zero, compete to influence a large population of agents. In particular, the case in which the transition and utility functions depend on the state, the action of the controllers, and the mean of the state and the actions, is investigated. The game is analyzed and explicit expressions for the Nash equilibrium strategies are derived. Moreover, two policy optimization methods that rely on policy gradient are proposed for both model-based and sample-based frameworks. In the first case, the gradients are computed exactly using the model whereas they are estimated using Monte-Carlo simulations in the second case. Numerical experiments show the convergence of the two players' controls as well as the utility function when the two algorithms are used in different scenarios.

📄 PDF Abstract BibTeX arXiv:2009.02146

Code (0)

등록된 구현이 없습니다.

Tasks

Vocal Bursts Type Prediction

Similar Papers 제목 키워드 기반

Linear-Quadratic Zero-Sum Mean-Field Type Games: Optimality Conditions and Policy Optimization

2020-09-01 · René Carmona, Kenza Hamidouche, Mathieu Laurière, Zongjun Tan

In this paper, zero-sum mean-field type games (ZSMFTG) with linear dynamics and quadratic cost are studied under infinite-horizon discounted utility function. ZSMFTG are a class of games in which two decision makers whos…

Policy Optimization Provably Converges to Nash Equilibria in Zero-Sum Linear Quadratic Games

2019-05-31 · NeurIPS 2019 12 · Kaiqing Zhang, Zhuoran Yang, Tamer Başar

We study the global convergence of policy optimization for finding the Nash equilibria (NE) in zero-sum linear quadratic (LQ) games. To this end, we first investigate the landscape of LQ games, viewing it as a nonconvex-…

Reinforcement Learning

Distributed Reinforcement Learning for Decentralized Linear Quadratic Control: A Derivative-Free Policy Optimization Approach

2019-12-19 · L4DC 2020 6 · Ying-Ying Li, Yujie Tang, Runyu Zhang, Na Li

This paper considers a distributed reinforcement learning problem for decentralized linear quadratic control with partial state observations and local costs. We propose a Zero-Order Distributed Policy Optimization algori…

Reinforcement LearningReinforcement Learning (RL)

Derivative-Free Methods for Policy Optimization: Guarantees for Linear Quadratic Systems

2018-12-20 · Dhruv Malik, Ashwin Pananjady, Kush Bhatia, Koulik Khamaru 외

We study derivative-free methods for policy optimization over the class of linear policies. We focus on characterizing the convergence rate of these methods when applied to linear-quadratic systems, and study various set…

Policy Optimization for Markovian Jump Linear Quadratic Control: Gradient-Based Methods and Global Convergence

2020-11-24 · Joao Paulo Jansch-Porto, Bin Hu, Geir Dullerud

Recently, policy optimization for control purposes has received renewed attention due to the increasing interest in reinforcement learning. In this paper, we investigate the global convergence of gradient-based policy op…

Policy Gradient Methods