Delta family approach for the stochastic control problems of utility maximization
In this paper, we propose a new approach for stochastic control problems arising from utility maximization. The main idea is to directly start from the dynamical programming equation and compute the conditional expectation using a novel representation of the conditional density function through the Dirac Delta function and the corresponding series representation. We obtain an explicit series representation of the value function, whose coefficients are expressed through integration of the value function at a later time point against a chosen basis function. Thus we are able to set up a recursive integration time-stepping scheme to compute the optimal value function given the known terminal condition, e.g. utility function. Due to tensor decomposition property of the Dirac Delta function in high dimensions, it is straightforward to extend our approach to solving high-dimensional stochastic control problems. The backward recursive nature of the method also allows for solving stochastic control and stopping problems, i.e. mixed control problems. We illustrate the method through solving some two-dimensional stochastic control (and stopping) problems, including the case under the classical and rough Heston stochastic volatility models, and stochastic local volatility models such as the stochastic alpha beta rho (SABR) model.
Code (0)
등록된 구현이 없습니다.
Tasks
Tensor DecompositionSimilar Papers 제목 키워드 기반
An interior-point stochastic approximation method and an L1-regularized delta rule
The stochastic approximation method is behind the solution to many important, actively-studied problems in machine learning. Despite its far-reaching application, there is almost no work on applying stochastic approximat…
BIG-bench Machine Learningfeature selectionEfficient and Adaptive Posterior Sampling Algorithms for Bandits
We study Thompson Sampling-based algorithms for stochastic bandits with bounded rewards. As the existing problem-dependent regret bound for Thompson Sampling with Gaussian priors [Agrawal and Goyal, 2017] is vacuous when…
Thompson SamplingRobust Data Valuation with Weighted Banzhaf Values
Data valuation, a principled way to rank the importance of each training datum, has become increasingly important. However, existing value-based approaches (e.g., Shapley) are known to suffer from the stochasticity inher…
Time-inconsistent contract theory
This paper investigates the moral hazard problem in finite horizon with both continuous and lump-sum payments, involving a time-inconsistent sophisticated agent and a standard utility maximiser principal. Building upon t…
Random Shuffling and Resets for the Non-stationary Stochastic Bandit Problem
We consider a non-stationary formulation of the stochastic multi-armed bandit where the rewards are no longer assumed to be identically distributed. For the best-arm identification task, we introduce a version of Success…