Optimizing Quantiles in Preference-based Markov Decision Processes
In the Markov decision process model, policies are usually evaluated by expected cumulative rewards. As this decision criterion is not always suitable, we propose in this paper an algorithm for computing a policy optimal for the quantile criterion. Both finite and infinite horizons are considered. Finally we experimentally evaluate our approach on random MDPs and on a data center control problem.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Quantile Markov Decision Process
The goal of a traditional Markov decision process (MDP) is to maximize expected cumulativereward over a defined horizon (possibly infinite). In many applications, however, a decision maker may beinterested in optimizing …
Solving Multi-Objective MDP with Lexicographic Preference: An application to stochastic planning with multiple quantile objective
In most common settings of Markov Decision Process (MDP), an agent evaluate a policy based on expectation of (discounted) sum of rewards. However in many applications this criterion might not be suitable from two perspec…
Autonomous DrivingDual theory of choice with multivariate risks
We propose a multivariate extension of Yaari's dual theory of choice under risk. We show that a decision maker with a preference relation on multidimensional prospects that preserves first order stochastic dominance and …
AttributeFair Resource Allocation in Weakly Coupled Markov Decision Processes
We consider fair resource allocation in sequential decision-making environments modeled as weakly coupled Markov decision processes, where resource constraints couple the action spaces of $N$ sub-Markov decision processe…
Decision MakingDeep Reinforcement LearningFairnessSequential Decision MakingSafe Reinforcement Learning in Constrained Markov Decision Processes
Safe reinforcement learning has been a promising approach for optimizing the policy of an agent that operates in safety-critical applications. In this paper, we propose an algorithm, SNO-MDP, that explores and optimizes …
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Safe Reinforcement Learning