paper-with-me

홈 › Papers

Refined Analysis of FPL for Adversarial Markov Decision Processes

2020-08-21 · Yuanhao Wang, Kefan Dong

We consider the adversarial Markov Decision Process (MDP) problem, where the rewards for the MDP can be adversarially chosen, and the transition function can be either known or unknown. In both settings, Follow-the-PerturbedLeader (FPL) based algorithms have been proposed in previous literature. However, the established regret bounds for FPL based algorithms are worse than algorithms based on mirrordescent. We improve the analysis of FPL based algorithms in both settings, matching the current best regret bounds using faster and simpler algorithms.

📄 PDF Abstract BibTeX arXiv:2008.09251

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bandit Linear Optimization for Sequential Decision Making and Extensive-Form Games

2021-03-08 · Gabriele Farina, Robin Schmucker, Tuomas Sandholm

Tree-form sequential decision making (TFSDM) extends classical one-shot decision making by modeling tree-form interactions between an agent and a potentially adversarial environment. It captures the online decision-makin…

counterfactualDecision MakingFormSequential Decision Making

Best-Effort Policies for Robust Markov Decision Processes

2025-08-11 · Alessandro Abate, Thom Badings, Giuseppe De Giacomo, Francesco Fabiano arxiv

We study the common generalization of Markov decision processes (MDPs) with sets of transition probabilities, known as robust MDPs (RMDPs). A standard goal in RMDPs is to compute a policy that maximizes the expected retu…

Reinforcement Learning in Robust Markov Decision Processes

2013-12-01 · NeurIPS 2013 12 · Shiau Hong Lim, Huan Xu, Shie Mannor

An important challenge in Markov decision processes is to ensure robustness with respect to unexpected or adversarial system behavior while taking advantage of well-behaving parts of the system. We consider a problem set…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Narrowing the Gap between Adversarial and Stochastic MDPs via Policy Optimization

2024-07-08 · Daniil Tiapkin, Evgenii Chzhen, Gilles Stoltz

In this paper, we consider the problem of learning in adversarial Markov decision processes [MDPs] with an oblivious adversary in a full-information setting. The agent interacts with an environment during $T$ episodes, e…

Online Convex Optimization in Adversarial Markov Decision Processes

2019-05-19 · Aviv Rosenberg, Yishay Mansour

We consider online learning in episodic loop-free Markov decision processes (MDPs), where the loss function can change arbitrarily between episodes, and the transition function is not known to the learner. We show $\tild…