paper-with-me

홈 › Papers

Detecting Spiky Corruption in Markov Decision Processes

2019-06-30 · Jason Mancuso, Tomasz Kisielewski, David Lindner, Alok Singh

Current reinforcement learning methods fail if the reward function is imperfect, i.e. if the agent observes reward different from what it actually receives. We study this problem within the formalism of Corrupt Reward Markov Decision Processes (CRMDPs). We show that if the reward corruption in a CRMDP is sufficiently "spiky", the environment is solvable. We fully characterize the regret bound of a Spiky CRMDP, and introduce an algorithm that is able to detect its corrupt states. We show that this algorithm can be used to learn the optimal policy with any common reinforcement learning algorithm. Finally, we investigate our algorithm in a pair of simple gridworld environments, finding that our algorithm can detect the corrupt states and learn the optimal policy despite the corruption.

📄 PDF Abstract BibTeX arXiv:1907.00452

Code (1)

jvmancuso/safe-grid-agents 공식 구현 tf

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Corruption-Robust Offline Reinforcement Learning with General Function Approximation

2023-10-23 · NeurIPS 2023 11 · Chenlu Ye, Rui Yang, Quanquan Gu, Tong Zhang

We investigate the problem of corruption robustness in offline reinforcement learning (RL) with general function approximation, where an adversary can corrupt each sample in the offline dataset, and the corruption level …

Offline RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Learning Constrained Markov Decision Processes With Non-stationary Rewards and Constraints

2024-05-23 · Francesco Emanuele Stradi, Anna Lunghi, Matteo Castiglioni, Alberto Marchesi 외

In constrained Markov decision processes (CMDPs) with adversarial rewards and constraints, a well-known impossibility result prevents any algorithm from attaining both sublinear regret and sublinear constraint violation,…

Corruption-Robust Algorithms with Uncertainty Weighting for Nonlinear Contextual Bandits and Markov Decision Processes

2022-12-12 · Chenlu Ye, Wei Xiong, Quanquan Gu, Tong Zhang

Despite the significant interest and progress in reinforcement learning (RL) problems with adversarial corruption, current works are either confined to the linear setting or lead to an undesired $\tilde{O}(\sqrt{T}\zeta)…

Multi-Armed BanditsReinforcement Learning (RL)

Lattice investment projects support process model with corruption

2019-01-25

Lattice investment projects support process model with corruption is formulated and analyzed. The model is based on the Ising lattice model of ferromagnetic but takes deal with the social phenomenon. Set of corruption ag…

Prehistory

A Dynamic Watermarking Algorithm for Finite Markov Decision Problems

2021-11-09 · Jiacheng Tang, Jiguo Song, Abhishek Gupta

Dynamic watermarking, as an active intrusion detection technique, can potentially detect replay attacks, spoofing attacks, and deception attacks in the feedback channel for control systems. In this paper, we develop a no…

Intrusion DetectionSensitivity