paper-with-me

Papers

Optimizing Quantiles in Preference-based Markov Decision Processes

2016-12-01 · Hugo Gilbert, Paul Weng, Yan Xu

In the Markov decision process model, policies are usually evaluated by expected cumulative rewards. As this decision criterion is not always suitable, we propose in this paper an algorithm for computing a policy optimal for the quantile criterion. Both finite and infinite horizons are considered. Finally we experimentally evaluate our approach on random MDPs and on a data center control problem.

📄 PDF Abstract BibTeX arXiv:1612.00094

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Quantile Markov Decision Process

2017-11-15 · Xiaocheng Li, Huaiyang Zhong, Margaret L. Brandeau

The goal of a traditional Markov decision process (MDP) is to maximize expected cumulativereward over a defined horizon (possibly infinite). In many applications, however, a decision maker may beinterested in optimizing …

Solving Multi-Objective MDP with Lexicographic Preference: An application to stochastic planning with multiple quantile objective

2017-05-10 · Yan Li, Zhaohan Sun

In most common settings of Markov Decision Process (MDP), an agent evaluate a policy based on expectation of (discounted) sum of rewards. However in many applications this criterion might not be suitable from two perspec…

Autonomous Driving

Dual theory of choice with multivariate risks

2021-02-04 · Alfred Galichon, Marc Henry

We propose a multivariate extension of Yaari's dual theory of choice under risk. We show that a decision maker with a preference relation on multidimensional prospects that preserves first order stochastic dominance and …

Attribute

Fair Resource Allocation in Weakly Coupled Markov Decision Processes

2024-11-14 · Xiaohui Tu, Yossiri Adulyasak, Nima Akbarzadeh, Erick Delage

We consider fair resource allocation in sequential decision-making environments modeled as weakly coupled Markov decision processes, where resource constraints couple the action spaces of $N$ sub-Markov decision processe…

Decision MakingDeep Reinforcement LearningFairnessSequential Decision Making

Safe Reinforcement Learning in Constrained Markov Decision Processes

2020-08-15 · ICML 2020 1 · Akifumi Wachi, Yanan Sui

Safe reinforcement learning has been a promising approach for optimizing the policy of an agent that operates in safety-critical applications. In this paper, we propose an algorithm, SNO-MDP, that explores and optimizes …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Safe Reinforcement Learning