Continuous-Time Mean-Variance Portfolio Selection: A Reinforcement Learning Framework
We approach the continuous-time mean-variance (MV) portfolio selection with reinforcement learning (RL). The problem is to achieve the best tradeoff between exploration and exploitation, and is formulated as an entropy-regularized, relaxed stochastic control problem. We prove that the optimal feedback policy for this problem must be Gaussian, with time-decaying variance. We then establish connections between the entropy-regularized MV and the classical MV, including the solvability equivalence and the convergence as exploration weighting parameter decays to zero. Finally, we prove a policy improvement theorem, based on which we devise an implementable RL algorithm. We find that our algorithm outperforms both an adaptive control based method and a deep neural networks based algorithm by a large margin in our simulations.
Code (1)
Tasks
Continuous ControlPortfolio Optimizationreinforcement-learningReinforcement LearningReinforcement Learning (RL)Similar Papers 제목 키워드 기반
Continuous-Time Mean-Variance Portfolio Selection with Constraints on Wealth and Portfolio
We consider continuous-time mean-variance portfolio selection with bankruptcy prohibition under convex cone portfolio constraints. This is a long-standing and difficult problem not only because of its theoretical signifi…
Continuous-Time Path-Dependent Exploratory Mean-Variance Portfolio Construction
In this paper, we present an extended exploratory continuous-time mean-variance framework for portfolio management. Our strategy involves a new clustering method based on simulated annealing, which allows for more practi…
ClusteringManagementExplicit solutions for continuous time mean-variance portfolio selection with nonlinear wealth equations
This paper concerns the continuous time mean-variance portfolio selection problem with a special nonlinear wealth equation. This nonlinear wealth equation has a nonsmooth coefficient and the dual method developed in [6] …
The Exploratory Multi-Asset Mean-Variance Portfolio Selection using Reinforcement Learning
In this paper, we study the continuous-time multi-asset mean-variance (MV) portfolio selection using a reinforcement learning (RL) algorithm, specifically the soft actor-critic (SAC) algorithm, in the time-varying financ…
Reinforcement Learning (RL)Mean-variance portfolio selection with tracking error penalization
This paper studies a variation of the continuous-time mean-variance portfolio selection where a tracking-error penalization is added to the mean-variance criterion. The tracking error term penalizes the distance between …