Regularizing a Model-based Policy Stationary Distribution to Stabilize Offline Reinforcement Learning
Offline reinforcement learning (RL) extends the paradigm of classical RL algorithms to purely learning from static datasets, without interacting with the underlying environment during the learning process. A key challenge of offline RL is the instability of policy training, caused by the mismatch between the distribution of the offline data and the undiscounted stationary state-action distribution of the learned policy. To avoid the detrimental impact of distribution mismatch, we regularize the undiscounted stationary distribution of the current policy towards the offline data during the policy optimization process. Further, we train a dynamics model to both implement this regularization and better estimate the stationary distribution of the current policy, reducing the error induced by distribution mismatch. On a wide range of continuous-control offline RL datasets, our method indicates competitive performance, which validates our algorithm. The code is publicly available.
Code (1)
Tasks
continuous-controlContinuous ControlOffline RLreinforcement-learningReinforcement Learning (RL)Similar Papers 제목 키워드 기반
Offline Imitation Learning with Suboptimal Demonstrations via Relaxed Distribution Matching
Offline imitation learning (IL) promises the ability to learn performant policies from pre-collected demonstrations without interactions with the environment. However, imitating behaviors fully offline typically requires…
continuous-controlContinuous ControlImitation LearningOut-of-Distribution Adaptation in Offline RL: Counterfactual Reasoning via Causal Normalizing Flows
Despite notable successes of Reinforcement Learning (RL), the prevalent use of an online learning paradigm prevents its widespread adoption, especially in hazardous or costly scenarios. Offline RL has emerged as an alter…
Causal InferencecounterfactualCounterfactual ReasoningOffline RL+2COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation
We consider the offline constrained reinforcement learning (RL) problem, in which the agent aims to compute a policy that maximizes expected return while satisfying given cost constraints, learning only from a pre-collec…
Offline RLOff-policy evaluationreinforcement-learningReinforcement Learning+1OptiDICE: Offline Policy Optimization via Stationary Distribution Correction Estimation
We consider the offline reinforcement learning (RL) setting where the agent aims to optimize the policy solely from the data without further environment interactions. In offline RL, the distributional shift becomes the p…
Offline RLReinforcement Learning (RL)DemoDICE: Offline Imitation Learning with Supplementary Imperfect Demonstrations
We consider offline imitation learning (IL), which aims to mimic the expert's behavior from its demonstration without further interaction with the environment. One of the main challenges in offline IL is to deal with th…
Imitation Learning