paper-with-me

Papers

Robust Reinforcement Learning for Continuous Control with Model Misspecification

2019-06-18 · ICLR 2020 1 · Daniel J. Mankowitz, Nir Levine, Rae Jeong, Yuanyuan Shi, Jackie Kay, Abbas Abdolmaleki, Jost Tobias Springenberg, Timothy Mann, Todd Hester, Martin Riedmiller

We provide a framework for incorporating robustness -- to perturbations in the transition dynamics which we refer to as model misspecification -- into continuous control Reinforcement Learning (RL) algorithms. We specifically focus on incorporating robustness into a state-of-the-art continuous control RL algorithm called Maximum a-posteriori Policy Optimization (MPO). We achieve this by learning a policy that optimizes for a worst case expected return objective and derive a corresponding robust entropy-regularized Bellman contraction operator. In addition, we introduce a less conservative, soft-robust, entropy-regularized objective with a corresponding Bellman operator. We show that both, robust and soft-robust policies, outperform their non-robust counterparts in nine Mujoco domains with environment perturbations. In addition, we show improved robust performance on a high-dimensional, simulated, dexterous robotic hand. Finally, we present multiple investigative experiments that provide a deeper insight into the robustness framework. This includes an adaptation to another continuous control RL algorithm as well as learning the uncertainty set from offline data. Performance videos can be found online at https://sites.google.com/view/robust-rl.

📄 PDF Abstract BibTeX arXiv:1906.07516

Code (0)

등록된 구현이 없습니다.

Tasks

continuous-controlContinuous ControlMuJoCoreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Robust Constrained Reinforcement Learning for Continuous Control with Model Misspecification

2020-10-20 · Daniel J. Mankowitz, Dan A. Calian, Rae Jeong, Cosmin Paduraru 외

Many real-world physical control systems are required to satisfy constraints upon deployment. Furthermore, real-world systems are often subject to effects such as non-stationarity, wear-and-tear, uncalibrated sensors and…

continuous-controlContinuous ControlMuJoCoreinforcement-learning+2

Robust Reinforcement Learning under model misspecification

2021-03-29 · Lebin Yu, Jian Wang, Xudong Zhang

Reinforcement learning has achieved remarkable performance in a wide range of tasks these days. Nevertheless, some unsolved problems limit its applications in real-world control. One of them is model misspecification, a …

Adversarial Attackmodelreinforcement-learningReinforcement Learning+1

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification

2026-06-04 · Haoyang Hong, Zichen Wang, Quanquan Gu, Huazheng Wang arxiv

We study KL-regularized contextual bandits and episodic reinforcement learning (RL) under general function approximation with model misspecification. Existing guarantees rely on realizability and therefore do not extend …

Reinforcement Learning

Game-Theoretic Robust Reinforcement Learning Handles Temporally-Coupled Perturbations

2023-07-22 · Yongyuan Liang, Yanchao Sun, Ruijie Zheng, Xiangyu Liu 외

Deploying reinforcement learning (RL) systems requires robustness to uncertainty and model misspecification, yet prior robust RL methods typically only study noise introduced independently across time. However, practical…

continuous-controlContinuous Controlreinforcement-learningReinforcement Learning+1

Double robust inference for continuous updating GMM

2021-05-18 · Frank Kleibergen, Zhaoguo Zhan

We propose the double robust Lagrange multiplier (DRLM) statistic for testing hypotheses specified on the pseudo-true value of the structural parameters in the generalized method of moments. The pseudo-true value is defi…