paper-with-me

Papers

DCE: Offline Reinforcement Learning With Double Conservative Estimates

2022-09-27 · Chen Zhao, Kai Xing Huang, Chun Yuan

Offline Reinforcement Learning has attracted much interest in solving the application challenge for traditional reinforcement learning. Offline reinforcement learning uses previously-collected datasets to train agents without any interaction. For addressing the overestimation of OOD (out-of-distribution) actions, conservative estimates give a low value for all inputs. Previous conservative estimation methods are usually difficult to avoid the impact of OOD actions on Q-value estimates. In addition, these algorithms usually need to lose some computational efficiency to achieve the purpose of conservative estimation. In this paper, we propose a simple conservative estimation method, double conservative estimates (DCE), which use two conservative estimation method to constraint policy. Our algorithm introduces V-function to avoid the error of in-distribution action while implicit achieving conservative estimation. In addition, our algorithm uses a controllable penalty term changing the degree of conservatism in training. We theoretically show how this method influences the estimation of OOD actions and in-distribution actions. Our experiment separately shows that two conservative estimation methods impact the estimation of all state-action. DCE demonstrates the state-of-the-art performance on D4RL.

📄 PDF Abstract BibTeX arXiv:2209.13132

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyD4RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Exploiting Generalization in Offline Reinforcement Learning via Unseen State Augmentations

2023-08-07 · Nirbhay Modhe, Qiaozi Gao, Ashwin Kalyan, Dhruv Batra 외

Offline reinforcement learning (RL) methods strike a balance between exploration and exploitation by conservative value estimation -- penalizing values of unseen states and actions. Model-free methods penalize values at …

Offline RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)

CROP: Conservative Reward for Model-based Offline Policy Optimization

2023-10-26 · Hao Li, Xiao-Hu Zhou, Xiao-Liang Xie, Shi-Qi Liu 외

Offline reinforcement learning (RL) aims to optimize policy using collected data without online interactions. Model-based approaches are particularly appealing for addressing offline RL challenges due to their capability…

D4RLOffline RLReinforcement Learning (RL)

Multi-objective Optimization of Notifications Using Offline Reinforcement Learning

2022-07-07 · Prakruthi Prabhakar, Yiping Yuan, Guangyu Yang, Wensheng Sun 외

Mobile notification systems play a major role in a variety of applications to communicate, send alerts and reminders to the users to inform them about news, events or messages. In this paper, we formulate the near-real-t…

Q-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Strategically Conservative Q-Learning

2024-06-06 · Yutaka Shimizu, Joey Hong, Sergey Levine, Masayoshi Tomizuka

Offline reinforcement learning (RL) is a compelling paradigm to extend RL's practical utility by leveraging pre-collected, static datasets, thereby avoiding the limitations associated with collecting online interactions.…

D4RLOffline RLQ-LearningReinforcement Learning (RL)

Conservative Bayesian Model-Based Value Expansion for Offline Policy Optimization

2022-10-07 · Jihwan Jeong, Xiaoyu Wang, Michael Gimelfarb, Hyunwoo Kim 외

Offline reinforcement learning (RL) addresses the problem of learning a performant policy from a fixed batch of data collected by following some behavior policy. Model-based approaches are particularly appealing in the o…

continuous-controlContinuous ControlD4RLReinforcement Learning (RL)