Markowitz Meets Bellman: Knowledge-distilled Reinforcement Learning for Portfolio Management
Investment portfolios, central to finance, balance potential returns and risks. This paper introduces a hybrid approach combining Markowitz's portfolio theory with reinforcement learning, utilizing knowledge distillation for training agents. In particular, our proposed method, called KDD (Knowledge Distillation DDPG), consist of two training stages: supervised and reinforcement learning stages. The trained agents optimize portfolio assembly. A comparative analysis against standard financial models and AI frameworks, using metrics like returns, the Sharpe ratio, and nine evaluation indices, reveals our model's superiority. It notably achieves the highest yield and Sharpe ratio of 2.03, ensuring top profitability with the lowest risk in comparable return scenarios.
Code (0)
등록된 구현이 없습니다.
Tasks
Knowledge DistillationManagementreinforcement-learningReinforcement LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Continuous-Time Portfolio Choice Under Monotone Mean-Variance Preferences-Stochastic Factor Case
We consider an incomplete market with a nontradable stochastic factor and a continuous time investment problem with an optimality criterion based on monotone mean-variance preferences. We formulate it as a stochastic dif…
Bellman Meets Hawkes: Model-Based Reinforcement Learning via Temporal Point Processes
We consider a sequential decision making problem where the agent faces the environment characterized by the stochastic discrete events and seeks an optimal intervention policy such that its long-term reward is maximized.…
Decision MakingModel-based Reinforcement LearningPoint Processesreinforcement-learning+3Robust Markowitz mean-variance portfolio selection under ambiguous covariance matrix *
This paper studies a robust continuous-time Markowitz portfolio selection pro\-blem where the model uncertainty carries on the covariance matrix of multiple risky assets. This problem is formulated into a min-max mean-va…
Continuous-time Markowitz's mean-variance model under different borrowing and saving rates
We study Markowitz's mean-variance portfolio selection problem in a continuous-time Black-Scholes market with different borrowing and saving rates. The associated Hamilton-Jacobi-Bellman equation is fully nonlinear. Usin…
Distilled Reinforcement Learning for LLM Post-training
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). Howeve…
Reinforcement Learning