Off-Policy Optimization of Portfolio Allocation Policies under Constraints
The dynamic portfolio optimization problem in finance frequently requires learning policies that adhere to various constraints, driven by investor preferences and risk. We motivate this problem of finding an allocation policy within a sequential decision making framework and study the effects of: (a) using data collected under previously employed policies, which may be sub-optimal and constraint-violating, and (b) imposing desired constraints while computing near-optimal policies with this data. Our framework relies on solving a minimax objective, where one player evaluates policies via off-policy estimators, and the opponent uses an online learning strategy to control constraint violations. We extensively investigate various choices for off-policy estimation and their corresponding optimization sub-routines, and quantify their impact on computing constraint-aware allocation policies. Our study shows promising results for constructing such policies when back-tested on historical equities data, under various regimes of operation, dimensionality and constraints.
Code (1)
Tasks
Decision MakingPortfolio OptimizationSequential Decision MakingSimilar Papers 제목 키워드 기반
Deep Stock Trading: A Hierarchical Reinforcement Learning Framework for Portfolio Optimization and Order Execution
Portfolio management via reinforcement learning is at the forefront of fintech research, which explores how to optimally reallocate a fund into different financial assets over the long term by trial-and-error. Existing m…
Hierarchical Reinforcement LearningManagementPortfolio OptimizationReinforcement Learning (RL)MetaTrader: An Reinforcement Learning Approach Integrating Diverse Policies for Portfolio Optimization
Portfolio management is a fundamental problem in finance. It involves periodic reallocations of assets to maximize the expected returns within an appropriate level of risk exposure. Deep reinforcement learning (RL) has b…
Decision MakingDeep Reinforcement LearningImitation LearningManagement+4Plan Before You Trade: Inference-Time Optimization for RL Trading Agents
Reinforcement learning agents for portfolio management are typically trained and deployed as static policies, with no mechanism for using price forecasts at inference time. We propose $\text{FPILOT}$ (**Fin**ancial **P**…
Reinforcement LearningDynamic Factor Model-Based Multiperiod Mean-Variance Portfolio Selection with Portfolio Constraints
Motivated by practical applications, we explore the constrained multi-period mean-variance portfolio selection problem within a market characterized by a dynamic factor model. This model captures predictability in asset …
Onflow: an online portfolio allocation algorithm
We introduce Onflow, a reinforcement learning technique that enables online optimization of portfolio allocation policies based on gradient flows. We devise dynamic allocations of an investment portfolio to maximize its …
Stochastic Optimization