paper-with-me

홈 › Papers

On the Expressivity of Multidimensional Markov Reward

2023-07-22 · Shuwa Miura

We consider the expressivity of Markov rewards in sequential decision making under uncertainty. We view reward functions in Markov Decision Processes (MDPs) as a means to characterize desired behaviors of agents. Assuming desired behaviors are specified as a set of acceptable policies, we investigate if there exists a scalar or multidimensional Markov reward function that makes the policies in the set more desirable than the other policies. Our main result states both necessary and sufficient conditions for the existence of such reward functions. We also show that for every non-degenerate set of deterministic policies, there exists a multidimensional Markov reward function that characterizes it

📄 PDF Abstract BibTeX arXiv:2307.12184

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingDecision Making Under UncertaintySequential Decision Making

Similar Papers 제목 키워드 기반

On the Expressivity of Markov Reward

2021-11-01 · NeurIPS 2021 12 · David Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho 외

Reward is the driving force for reinforcement-learning agents. This paper is dedicated to understanding the expressivity of reward as a way to capture tasks that we would want an agent to perform. We frame this study aro…

On The Expressivity of Objective-Specification Formalisms in Reinforcement Learning

2023-10-18 · Rohan Subramani, Marcus Williams, Max Heitmann, Halfdan Holm 외

Most algorithms in reinforcement learning (RL) require that the objective is formalised with a Markovian reward function. However, it is well-known that certain tasks cannot be expressed by means of an objective in the M…

Multi-Objective Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

On the Limitations of Markovian Rewards to Express Multi-Objective, Risk-Sensitive, and Modal Tasks

2024-01-26 · Joar Skalse, Alessandro Abate

In this paper, we study the expressivity of scalar, Markovian reward functions in Reinforcement Learning (RL), and identify several limitations to what they can express. Specifically, we look at three classes of RL tasks…

Reinforcement Learning (RL)

Expressive Reward Synthesis with the Runtime Monitoring Language

2025-10-17 · Daniel Donnelly, Angelo Ferrando, Francesco Belardinelli arxiv

A key challenge in reinforcement learning (RL) is reward (mis)specification, whereby imprecisely defined reward functions can result in unintended, possibly harmful, behaviours. Indeed, reward functions in RL are typical…

Reinforcement Learning

Policy Synthesis and Reinforcement Learning for Discounted LTL

2023-05-26 · Rajeev Alur, Osbert Bastani, Kishor Jothimurugan, Mateo Perez 외

The difficulty of manually specifying reward functions has led to an interest in using linear temporal logic (LTL) to express objectives for reinforcement learning (RL). However, LTL has the downside that it is sensitive…

PAC learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1