paper-with-me

홈 › Papers

Continuous MDP Homomorphisms and Homomorphic Policy Gradient

2022-09-15 · Sahand Rezaei-Shoshtari, Rosie Zhao, Prakash Panangaden, David Meger, Doina Precup

Abstraction has been widely studied as a way to improve the efficiency and generalization of reinforcement learning algorithms. In this paper, we study abstraction in the continuous-control setting. We extend the definition of MDP homomorphisms to encompass continuous actions in continuous state spaces. We derive a policy gradient theorem on the abstract MDP, which allows us to leverage approximate symmetries of the environment for policy optimization. Based on this theorem, we propose an actor-critic algorithm that is able to learn the policy and the MDP homomorphism map simultaneously, using the lax bisimulation metric. We demonstrate the effectiveness of our method on benchmark tasks in the DeepMind Control Suite. Our method's ability to utilize MDP homomorphisms for representation learning leads to improved performance when learning from pixel observations.

📄 PDF Abstract BibTeX arXiv:2209.07364

Code (1)

sahandrez/homomorphic_policy_gradient 공식 구현 pytorch

Tasks

continuous-controlContinuous ControlPolicy Gradient MethodsReinforcement Learning (RL)Representation Learning

Similar Papers 제목 키워드 기반

Delayed homomorphic reinforcement learning for environments with delayed feedback

2026-04-04 · Jongsoo Lee, Jangwon Kim, Soohee Han arxiv

Reinforcement learning in real-world systems often involves delayed feedback, which breaks the Markov assumption and impedes both learning and control. Canonical augmentation-based approaches cause state-space explosion,…

Reinforcement Learning

Policy Gradient Methods in the Presence of Symmetries and State Abstractions

2023-05-09 · Prakash Panangaden, Sahand Rezaei-Shoshtari, Rosie Zhao, David Meger 외

Reinforcement learning (RL) on high-dimensional and complex problems relies on abstraction for improved efficiency and generalization. In this paper, we study abstraction in the continuous-control setting, and extend the…

continuous-controlContinuous ControlPolicy Gradient MethodsReinforcement Learning (RL)+1

Online Abstraction with MDP Homomorphisms for Deep Learning

2018-11-30 · Ondrej Biza, Robert Platt

Abstraction of Markov Decision Processes is a useful tool for solving complex problems, as it can ignore unimportant aspects of an environment, simplifying the process of learning an optimal policy. In this paper, we pro…

Deep Learning

Bounding Performance Loss in Approximate MDP Homomorphisms

2008-12-01 · NeurIPS 2008 12 · Jonathan Taylor, Doina Precup, Prakash Panagaden

We define a metric for measuring behavior similarity between states in a Markov decision process (MDP), in which action similarity is taken into account. We show that the kernel of our metric corresponds exactly to the c…

A Note on Categories about Rough Sets

2022-05-19 · Y. R. Syau, E. B. Lin, C. J. Liau

Using the concepts of category and functor, we provide some insights and prove an intrinsic property of the category ${\bf AprS}$ of approximation spaces and relation-preserving functions, the category ${\bf RCls}$ of ro…

Attribute