Continuous MDP Homomorphisms and Homomorphic Policy Gradient
Abstraction has been widely studied as a way to improve the efficiency and generalization of reinforcement learning algorithms. In this paper, we study abstraction in the continuous-control setting. We extend the definition of MDP homomorphisms to encompass continuous actions in continuous state spaces. We derive a policy gradient theorem on the abstract MDP, which allows us to leverage approximate symmetries of the environment for policy optimization. Based on this theorem, we propose an actor-critic algorithm that is able to learn the policy and the MDP homomorphism map simultaneously, using the lax bisimulation metric. We demonstrate the effectiveness of our method on benchmark tasks in the DeepMind Control Suite. Our method's ability to utilize MDP homomorphisms for representation learning leads to improved performance when learning from pixel observations.
Code (1)
Tasks
continuous-controlContinuous ControlPolicy Gradient MethodsReinforcement Learning (RL)Representation LearningSimilar Papers 제목 키워드 기반
Delayed homomorphic reinforcement learning for environments with delayed feedback
Reinforcement learning in real-world systems often involves delayed feedback, which breaks the Markov assumption and impedes both learning and control. Canonical augmentation-based approaches cause state-space explosion,…
Reinforcement LearningPolicy Gradient Methods in the Presence of Symmetries and State Abstractions
Reinforcement learning (RL) on high-dimensional and complex problems relies on abstraction for improved efficiency and generalization. In this paper, we study abstraction in the continuous-control setting, and extend the…
continuous-controlContinuous ControlPolicy Gradient MethodsReinforcement Learning (RL)+1Online Abstraction with MDP Homomorphisms for Deep Learning
Abstraction of Markov Decision Processes is a useful tool for solving complex problems, as it can ignore unimportant aspects of an environment, simplifying the process of learning an optimal policy. In this paper, we pro…
Deep LearningBounding Performance Loss in Approximate MDP Homomorphisms
We define a metric for measuring behavior similarity between states in a Markov decision process (MDP), in which action similarity is taken into account. We show that the kernel of our metric corresponds exactly to the c…
A Note on Categories about Rough Sets
Using the concepts of category and functor, we provide some insights and prove an intrinsic property of the category ${\bf AprS}$ of approximation spaces and relation-preserving functions, the category ${\bf RCls}$ of ro…
Attribute