Cumulative Prospect Theory Meets Reinforcement Learning: Prediction and Control
Cumulative prospect theory (CPT) is known to model human decisions well, with substantial empirical evidence supporting this claim. CPT works by distorting probabilities and is more general than the classic expected utility and coherent risk measures. We bring this idea to a risk-sensitive reinforcement learning (RL) setting and design algorithms for both estimation and control. The RL setting presents two particular challenges when CPT is applied: estimating the CPT objective requires estimations of the entire distribution of the value function and finding a randomized optimal policy. The estimation scheme that we propose uses the empirical distribution to estimate the CPT-value of a random variable. We then use this scheme in the inner loop of a CPT-value optimization procedure that is based on the well-known simulation optimization idea of simultaneous perturbation stochastic approximation (SPSA). We provide theoretical convergence guarantees for all the proposed algorithms and also illustrate the usefulness of CPT-based criteria in a traffic signal control application.
Code (0)
등록된 구현이 없습니다.
Tasks
Predictionreinforcement-learningReinforcement LearningReinforcement Learning (RL)Traffic Signal ControlSimilar Papers 제목 키워드 기반
Risk-Sensitive Reinforcement Learning via Policy Gradient Search
The objective in a traditional reinforcement learning (RL) problem is to find a policy that optimizes the expected value of a performance metric such as the infinite-horizon cumulative discounted or long-run average cost…
Policy Gradient Methodsreinforcement-learningReinforcement LearningReinforcement Learning (RL)Optimal Investment with Transaction Costs under Cumulative Prospect Theory in Discrete Time
We study optimal investment problems under the framework of cumulative prospect theory (CPT). A CPT investor makes investment decisions in a single-period financial market with transaction costs. The objective is to seek…
Multi-period investment strategies under Cumulative Prospect Theory
In this article, inspired by Shi, et al. we investigate the optimal portfolio selection with one risk-free asset and one risky asset in a multiple period setting under cumulative prospect theory (CPT). Compared with thei…
SensitivityOption Pricing with Greed and Fear Factor: The Rational Finance Approach
We explain the main concepts of Prospect Theory and Cumulative Prospect Theory within the framework of rational dynamic asset pricing theory. We derive option pricing formulas when asset returns are altered with a genera…
Beyond Expected Returns: A Policy Gradient Algorithm for Cumulative Prospect Theoretic Reinforcement Learning
The widely used expected utility theory has been shown to be empirically inconsistent with human preferences in the psychology and behavioral economy literatures. Cumulative Prospect Theory (CPT) has been developed to fi…
Reinforcement Learning (RL)