Provably efficient RL with Rich Observations via Latent State Decoding
We study the exploration problem in episodic MDPs with rich observations generated from a small number of latent states. Under certain identifiability assumptions, we demonstrate how to estimate a mapping from the observations to latent states inductively through a sequence of regression and clustering steps -- where previously decoded latent states provide labels for later regression problems -- and use it to construct good exploration policies. We provide finite-sample guarantees on the quality of the learned state decoding function and exploration policies, and complement our theory with an empirical evaluation on a class of hard exploration problems. Our method exponentially improves over $Q$-learning with na\"ive exploration, even when $Q$-learning has cheating access to latent states.
Code (1)
Tasks
ClusteringQ-LearningregressionSimilar Papers 제목 키워드 기반
Provably Filtering Exogenous Distractors using Multistep Inverse Dynamics
Many real-world applications of reinforcement learning (RL) require the agent to deal with high-dimensional observations such as those generated from a megapixel camera. Prior work has addressed such problems with repres…
Reinforcement Learning (RL)Representation LearningProvable Rich Observation Reinforcement Learning with Combinatorial Latent States
We propose a novel setting for reinforcement learning that combines two common real-world difficulties: presence of observations (such as camera images) and factored states (such as location of objects). In our setting, …
Contrastive Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Provably Efficient Exploration for Reinforcement Learning Using Unsupervised Learning
Motivated by the prevailing paradigm of using unsupervised learning for efficient exploration in reinforcement learning (RL) problems [tang2017exploration,bellemare2016unifying], we investigate when this paradigm is prov…
Efficient Explorationreinforcement-learningReinforcement LearningReinforcement Learning (RL)Provable RL with Exogenous Distractors via Multistep Inverse Dynamics
Many real-world applications of reinforcement learning (RL) require the agent to deal with high-dimensional observations such as those generated from a megapixel camera. Prior work has addressed such problems with repres…
Reinforcement Learning (RL)Representation LearningRich-Observation Reinforcement Learning with Continuous Latent Dynamics
Sample-efficiency and reliability remain major bottlenecks toward wide adoption of reinforcement learning algorithms in continuous settings with high-dimensional perceptual inputs. Toward addressing these challenges, we …
reinforcement-learningReinforcement LearningRepresentation Learning