paper-with-me

홈 › Papers

Reinforcement Learning in Rich-Observation MDPs using Spectral Methods

2016-11-11 · Kamyar Azizzadenesheli, Alessandro Lazaric, Animashree Anandkumar

Reinforcement learning (RL) in Markov decision processes (MDPs) with large state spaces is a challenging problem. The performance of standard RL algorithms degrades drastically with the dimensionality of state space. However, in practice, these large MDPs typically incorporate a latent or hidden low-dimensional structure. In this paper, we study the setting of rich-observation Markov decision processes (ROMDP), where there are a small number of hidden states which possess an injective mapping to the observation states. In other words, every observation state is generated through a single hidden state, and this mapping is unknown a priori. We introduce a spectral decomposition method that consistently learns this mapping, and more importantly, achieves it with low regret. The estimated mapping is integrated into an optimistic RL algorithm (UCRL), which operates on the estimated hidden space. We derive finite-time regret bounds for our algorithm with a weak dependence on the dimensionality of the observed space. In fact, our algorithm asymptotically achieves the same average regret as the oracle UCRL algorithm, which has the knowledge of the mapping from hidden to observed spaces. Thus, we derive an efficient spectral RL algorithm for ROMDPs.

📄 PDF Abstract BibTeX arXiv:1611.03907

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Reinforcement Learning of POMDPs using Spectral Methods

2016-02-25 · Kamyar Azizzadenesheli, Alessandro Lazaric, Animashree Anandkumar

We propose a new reinforcement learning algorithm for partially observable Markov decision processes (POMDP) based on spectral decomposition methods. While spectral methods have been previously employed for consistent le…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Experimental results : Reinforcement Learning of POMDPs using Spectral Methods

2017-05-07 · Kamyar Azizzadenesheli, Alessandro Lazaric, Animashree Anandkumar

We propose a new reinforcement learning algorithm for partially observable Markov decision processes (POMDP) based on spectral decomposition methods. While spectral methods have been previously employed for consistent le…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Efficient Reinforcement Learning in Block MDPs: A Model-free Representation Learning Approach

2022-01-31 · Xuezhou Zhang, Yuda Song, Masatoshi Uehara, Mengdi Wang 외

We present BRIEE (Block-structured Representation learning with Interleaved Explore Exploit), an algorithm for efficient reinforcement learning in Markov Decision Processes with block-structured dynamics (i.e., Block MDP…

reinforcement-learningReinforcement Learning (RL)Representation Learning

Sample-Efficient Reinforcement Learning of Undercomplete POMDPs

2020-06-22 · NeurIPS 2020 12 · Chi Jin, Sham M. Kakade, Akshay Krishnamurthy, Qinghua Liu

Partial observability is a common challenge in many reinforcement learning applications, which requires an agent to maintain memory, infer latent states, and integrate this past information into exploration. This challen…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

PAC Reinforcement Learning with Rich Observations

2016-02-08 · NeurIPS 2016 12 · Akshay Krishnamurthy, Alekh Agarwal, John Langford

We propose and study a new model for reinforcement learning with rich observations, generalizing contextual bandits to sequential decision making. These models require an agent to take actions based on observations (feat…

Decision MakingMulti-Armed Banditsreinforcement-learningReinforcement Learning+2