paper-with-me

Papers

Learning-Based Mean-Payoff Optimization in an Unknown MDP under Omega-Regular Constraints

2018-04-24 · Jan Křetínský, Guillermo A. Pérez, Jean-François Raskin

We formalize the problem of maximizing the mean-payoff value with high probability while satisfying a parity objective in a Markov decision process (MDP) with unknown probabilistic transition function and unknown reward function. Assuming the support of the unknown transition function and a lower bound on the minimal transition probability are known in advance, we show that in MDPs consisting of a single end component, two combinations of guarantees on the parity and mean-payoff objectives can be achieved depending on how much memory one is willing to use. (i) For all $\epsilon$ and $\gamma$ we can construct an online-learning finite-memory strategy that almost-surely satisfies the parity objective and which achieves an $\epsilon$-optimal mean payoff with probability at least $1 - \gamma$. (ii) Alternatively, for all $\epsilon$ and $\gamma$ there exists an online-learning infinite-memory strategy that satisfies the parity objective surely and which achieves an $\epsilon$-optimal mean payoff with probability at least $1 - \gamma$. We extend the above results to MDPs consisting of more than one end component in a natural way. Finally, we show that the aforementioned guarantees are tight, i.e. there are MDPs for which stronger combinations of the guarantees cannot be ensured.

📄 PDF Abstract BibTeX arXiv:1804.08924

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Omega-Regular Decision Processes

2023-12-14 · Ernst Moritz Hahn, Mateo Perez, Sven Schewe, Fabio Somenzi 외

Regular decision processes (RDPs) are a subclass of non-Markovian decision processes where the transition and reward functions are guarded by some regular property of the past (a lookback). While RDPs enable intuitive an…

Bayesian Optimization under Heavy-tailed Payoffs

2019-09-16 · NeurIPS 2019 12 · Sayak Ray Chowdhury, Aditya Gopalan

We consider black box optimization of an unknown function in the nonparametric Gaussian process setting when the noise in the observed function values can be heavy tailed. This is in contrast to existing literature that …

Bayesian Optimization

Average Reward Reinforcement Learning for Omega-Regular and Mean-Payoff Objectives

2025-05-21 · Milad Kazemi, Mateo Perez, Fabio Somenzi, Sadegh Soudjani 외

Recent advances in reinforcement learning (RL) have renewed focus on the design of reward functions that shape agent behavior. Manually designing reward functions is tedious and error-prone. A principled alternative is t…

Reinforcement Learning (RL)

Almost Optimal Algorithms for Linear Stochastic Bandits with Heavy-Tailed Payoffs

2018-10-25 · NeurIPS 2018 12 · Han Shao, Xiaotian Yu, Irwin King, Michael R. Lyu

In linear stochastic bandits, it is commonly assumed that payoffs are with sub-Gaussian noises. In this paper, under a weaker assumption on noises, we study the problem of \underline{lin}ear stochastic {\underline b}andi…

Low-rank Matrix Bandits with Heavy-tailed Rewards

2024-04-26 · Yue Kang, Cho-Jui Hsieh, Thomas C. M. Lee

In stochastic low-rank matrix bandit, the expected reward of an arm is equal to the inner product between its feature matrix and some unknown $d_1$ by $d_2$ low-rank parameter matrix $\Theta^*$ with rank $r \ll d_1\wedge…