paper-with-me

홈 › Papers

Robust Reinforcement Learning with Wasserstein Constraint

2020-06-01 · Linfang Hou, Liang Pang, Xin Hong, Yanyan Lan, Zhi-Ming Ma, Dawei Yin

Robust Reinforcement Learning aims to find the optimal policy with some extent of robustness to environmental dynamics. Existing learning algorithms usually enable the robustness through disturbing the current state or simulating environmental parameters in a heuristic way, which lack quantified robustness to the system dynamics (i.e. transition probability). To overcome this issue, we leverage Wasserstein distance to measure the disturbance to the reference transition kernel. With Wasserstein distance, we are able to connect transition kernel disturbance to the state disturbance, i.e. reduce an infinite-dimensional optimization problem to a finite-dimensional risk-aware problem. Through the derived risk-aware optimal Bellman equation, we show the existence of optimal robust policies, provide a sensitivity analysis for the perturbations, and then design a novel robust learning algorithm--Wasserstein Robust Advantage Actor-Critic algorithm (WRAAC). The effectiveness of the proposed algorithm is verified in the Cart-Pole environment.

📄 PDF Abstract BibTeX arXiv:2006.00945

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Wasserstein Robust Reinforcement Learning

2019-07-30 · Mohammed Amin Abdullah, Hang Ren, Haitham Bou Ammar, Vladimir Milenkovic 외

Reinforcement learning algorithms, though successful, tend to over-fit to training environments hampering their application to the real-world. This paper proposes $\text{W}\text{R}^{2}\text{L}$ -- a robust reinforcement …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

A Cramér Distance perspective on Quantile Regression based Distributional Reinforcement Learning

2021-10-01 · NeurIPS 2021 12 · Alix Lhéritier, Nicolas Bondoux

Distributional reinforcement learning (DRL) extends the value-based approach by approximating the full distribution over future returns instead of the mean only, providing a richer signal that leads to improved performan…

Distributional Reinforcement Learningquantile regressionregressionreinforcement-learning+1

A proof of imitation of Wasserstein inverse reinforcement learning for multi-objective optimization

2023-05-17 · Akira Kitaoka, Riki Eto

We prove Wasserstein inverse reinforcement learning enables the learner's reward values to imitate the expert's reward values in a finite iteration for multi-objective optimizations. Moreover, we prove Wasserstein invers…

reinforcement-learningReinforcement Learning

Reinforcement Learning with Wasserstein Distance Regularisation, with Applications to Multipolicy Learning

2018-02-12 · Mohammed Amin Abdullah, Aldo Pacchiano, Moez Draief

We describe an application of Wasserstein distance to Reinforcement Learning. The Wasserstein distance in question is between the distribution of mappings of trajectories of a policy into some metric space, and some othe…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Visual Transfer for Reinforcement Learning via Wasserstein Domain Confusion

2020-06-04 · Josh Roy, George Konidaris

We introduce Wasserstein Adversarial Proximal Policy Optimization (WAPPO), a novel algorithm for visual transfer in Reinforcement Learning that explicitly learns to align the distributions of extracted features between a…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)