paper-with-me

홈 › Papers

Regularizing a Model-based Policy Stationary Distribution to Stabilize Offline Reinforcement Learning

2022-06-14 · Shentao Yang, Yihao Feng, Shujian Zhang, Mingyuan Zhou

Offline reinforcement learning (RL) extends the paradigm of classical RL algorithms to purely learning from static datasets, without interacting with the underlying environment during the learning process. A key challenge of offline RL is the instability of policy training, caused by the mismatch between the distribution of the offline data and the undiscounted stationary state-action distribution of the learned policy. To avoid the detrimental impact of distribution mismatch, we regularize the undiscounted stationary distribution of the current policy towards the offline data during the policy optimization process. Further, we train a dynamics model to both implement this regularization and better estimate the stationary distribution of the current policy, reducing the error induced by distribution mismatch. On a wide range of continuous-control offline RL datasets, our method indicates competitive performance, which validates our algorithm. The code is publicly available.

📄 PDF Abstract BibTeX arXiv:2206.07166

Code (1)

shentao-yang/sdm-gan_icml2022 공식 구현 pytorch

Tasks

continuous-controlContinuous ControlOffline RLreinforcement-learningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Offline Imitation Learning with Suboptimal Demonstrations via Relaxed Distribution Matching

2023-03-05 · Lantao Yu, Tianhe Yu, Jiaming Song, Willie Neiswanger 외

Offline imitation learning (IL) promises the ability to learn performant policies from pre-collected demonstrations without interactions with the environment. However, imitating behaviors fully offline typically requires…

continuous-controlContinuous ControlImitation Learning

Out-of-Distribution Adaptation in Offline RL: Counterfactual Reasoning via Causal Normalizing Flows

2024-05-06 · Minjae Cho, Jonathan P. How, Chuangchuang Sun

Despite notable successes of Reinforcement Learning (RL), the prevalent use of an online learning paradigm prevents its widespread adoption, especially in hazardous or costly scenarios. Offline RL has emerged as an alter…

Causal InferencecounterfactualCounterfactual ReasoningOffline RL+2

COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation

2022-04-19 · ICLR 2022 4 · Jongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess 외

We consider the offline constrained reinforcement learning (RL) problem, in which the agent aims to compute a policy that maximizes expected return while satisfying given cost constraints, learning only from a pre-collec…

Offline RLOff-policy evaluationreinforcement-learningReinforcement Learning+1

OptiDICE: Offline Policy Optimization via Stationary Distribution Correction Estimation

2021-06-21 · Jongmin Lee, Wonseok Jeon, Byung-Jun Lee, Joelle Pineau 외

We consider the offline reinforcement learning (RL) setting where the agent aims to optimize the policy solely from the data without further environment interactions. In offline RL, the distributional shift becomes the p…

Offline RLReinforcement Learning (RL)

DemoDICE: Offline Imitation Learning with Supplementary Imperfect Demonstrations

2021-09-29 · ICLR 2022 4 · Geon-Hyeong Kim, Seokin Seo, Jongmin Lee, Wonseok Jeon 외

We consider offline imitation learning (IL), which aims to mimic the expert's behavior from its demonstration without further interaction with the environment. One of the main challenges in offline IL is to deal with th…

Imitation Learning