paper-with-me

홈 › Papers

Policy Gradient Methods in the Presence of Symmetries and State Abstractions

2023-05-09 · Prakash Panangaden, Sahand Rezaei-Shoshtari, Rosie Zhao, David Meger, Doina Precup

Reinforcement learning (RL) on high-dimensional and complex problems relies on abstraction for improved efficiency and generalization. In this paper, we study abstraction in the continuous-control setting, and extend the definition of Markov decision process (MDP) homomorphisms to the setting of continuous state and action spaces. We derive a policy gradient theorem on the abstract MDP for both stochastic and deterministic policies. Our policy gradient results allow for leveraging approximate symmetries of the environment for policy optimization. Based on these theorems, we propose a family of actor-critic algorithms that are able to learn the policy and the MDP homomorphism map simultaneously, using the lax bisimulation metric. Finally, we introduce a series of environments with continuous symmetries to further demonstrate the ability of our algorithm for action abstraction in the presence of such symmetries. We demonstrate the effectiveness of our method on our environments, as well as on challenging visual control tasks from the DeepMind Control Suite. Our method's ability to utilize MDP homomorphisms for representation learning leads to improved performance, and the visualizations of the latent space clearly demonstrate the structure of the learned abstraction.

📄 PDF Abstract BibTeX arXiv:2305.05666

Code (2)

sahandrez/homomorphic_policy_gradient 공식 구현 pytorch
sahandrez/rotate_suite 공식 구현

Tasks

continuous-controlContinuous ControlPolicy Gradient MethodsReinforcement Learning (RL)Representation Learning

Similar Papers 제목 키워드 기반

Continuous MDP Homomorphisms and Homomorphic Policy Gradient

2022-09-15 · Sahand Rezaei-Shoshtari, Rosie Zhao, Prakash Panangaden, David Meger 외

Abstraction has been widely studied as a way to improve the efficiency and generalization of reinforcement learning algorithms. In this paper, we study abstraction in the continuous-control setting. We extend the definit…

continuous-controlContinuous ControlPolicy Gradient MethodsReinforcement Learning (RL)+1

Modal Decomposition of the Linear Swing Equation in Networks with Symmetries

2021-07-22 · Kshitij Bhatta, Majeed Hayat, Francesco Sorrentino

Symmetries are widespread in physical, technological, biological, and social systems and networks, including power grids. The swing equation is a classic model for the dynamics of powergrid networks. The main goal of thi…

A Temporal-Difference Approach to Policy Gradient Estimation

2022-02-04 · Samuele Tosatto, Andrew Patterson, Martha White, A. Rupam Mahmood

The policy gradient theorem (Sutton et al., 2000) prescribes the usage of a cumulative discounted state distribution under the target policy to approximate the gradient. Most algorithms based on this theorem, in practice…

Riemannian Optimization for Hadamard Products of Low-Rank Matrices

2026-05-31 · Pratik Jawanpuria, Ankish Chandresh, Bamdev Mishra arxiv

The elementwise Hadamard product of two low-rank matrices provides a parameter-efficient model for data with multiplicative structure, but its modeling is challenging due to the presence of additional symmetries under co…

Multi-Agent MDP Homomorphic Networks

2021-10-09 · ICLR 2022 4 · Elise van der Pol, Herke van Hoof, Frans A. Oliehoek, Max Welling

This paper introduces Multi-Agent MDP Homomorphic Networks, a class of networks that allows distributed execution using only local information, yet is able to share experience between global symmetries in the joint state…