Large scale continuous-time mean-variance portfolio allocation via reinforcement learning
We propose to solve large scale Markowitz mean-variance (MV) portfolio allocation problem using reinforcement learning (RL). By adopting the recently developed continuous-time exploratory control framework, we formulate the exploratory MV problem in high dimensions. We further show the optimality of a multivariate Gaussian feedback policy, with time-decaying variance, in trading off exploration and exploitation. Based on a provable policy improvement theorem, we devise a scalable and data-efficient RL algorithm and conduct large scale empirical tests using data from the S&P 500 stocks. We found that our method consistently achieves over 10% annualized returns and it outperforms econometric methods and the deep RL method by large margins, for both long and medium terms of investment with monthly and daily trading.
Code (0)
등록된 구현이 없습니다.
Tasks
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Similar Papers 제목 키워드 기반
Bellman type strategy for the continuous time mean-variance model
To investigate a time-consistent optimal strategy for the continuous time mean-variance model, we develop a new method to establish the Bellman principle. Based on this new method, we obtain a time-consistent dynamic opt…
Vocal Bursts Type PredictionContinuous-Time Path-Dependent Exploratory Mean-Variance Portfolio Construction
In this paper, we present an extended exploratory continuous-time mean-variance framework for portfolio management. Our strategy involves a new clustering method based on simulated annealing, which allows for more practi…
ClusteringManagementEvolutionary game with stochastic payoffs in a finite island model
In this paper, we consider a two-player two-strategy game with random payoffs in a population subdivided into $d$ demes, each containing $N$ individuals at the beginning of any given generation and experiencing local ext…
Continuous-Time Mean-Variance Portfolio Selection: A Reinforcement Learning Framework
We approach the continuous-time mean-variance (MV) portfolio selection with reinforcement learning (RL). The problem is to achieve the best tradeoff between exploration and exploitation, and is formulated as an entropy-r…
Continuous ControlPortfolio Optimizationreinforcement-learningReinforcement Learning+1Continuous-Time Mean-Variance Portfolio Selection with Constraints on Wealth and Portfolio
We consider continuous-time mean-variance portfolio selection with bankruptcy prohibition under convex cone portfolio constraints. This is a long-standing and difficult problem not only because of its theoretical signifi…