Deterministic Policy Gradient for Learning Equilibrium in Time-Inconsistent Control Problems
In this paper, we develop a continuous-time model-free reinforcement learning algorithm to learn deterministic equilibrium policies in general time-inconsistent control problems. Utilizing the extended Hamilton-Jacobi-Bellman system, we recast the original time-inconsistent problem into an equivalent two-stage problem. In the first stage, for given auxiliary functions, we employ the deterministic policy gradient approach to learn an optimal policy in an auxiliary time-consistent control problem. In the second stage, given the updated policy, we exploit the inner fixed point iterations and some martingale characterizations to learn the auxiliary functions. As a theoretical contribution, we provide some mild model assumptions and establish the convergence of inner fixed point iterations. By repeating this actor-critic style of iterations across two stages, our algorithm aims to learn the equilibrium under different sources of time-inconsistency in a unified manner. The superior effectiveness of the proposed algorithm are illustrated in two classical financial applications with time-inconsistency: mean-variance portfolio management and optimal tracking portfolio under non-exponential discounting.
Code (0)
등록된 구현이 없습니다.
Tasks
Reinforcement LearningSimilar Papers 제목 키워드 기반
On the Time-Inconsistent Deterministic Linear-Quadratic Control
A fundamental theory of deterministic linear-quadratic (LQ) control is the equivalent relationship between control problems, two-point boundary value problems and Riccati equations. In this paper, we extend the equivalen…
Time-Inconsistent Stochastic Linear--Quadratic Control: Characterization and Uniqueness of Equilibrium
In this paper, we continue our study on a general time-inconsistent stochastic linear--quadratic (LQ) control problem originally formulated in [6]. We derive a necessary and sufficient condition for equilibrium controls …
Policy Gradient Methods Find the Nash Equilibrium in N-player General-sum Linear-quadratic Games
We consider a general-sum N-player linear-quadratic game with stochastic dynamics over a finite horizon and prove the global convergence of the natural policy gradient method to the Nash equilibrium. In order to prove th…
Policy Gradient MethodsExploratory Mean-Variance with Jumps: An Equilibrium Approach
Revisiting the continuous-time Mean-Variance (MV) Portfolio Optimization problem, we model the market dynamics with a jump-diffusion process and apply Reinforcement Learning (RL) techniques to facilitate informed explora…
Reinforcement LearningPortfolio OptimizationComputational Performance of Deep Reinforcement Learning to find Nash Equilibria
We test the performance of deep deterministic policy gradient (DDPG), a deep reinforcement learning algorithm, able to handle continuous state and action spaces, to learn Nash equilibria in a setting where firms compete …
Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)