paper-with-me

홈 › Papers

Efficient Policy Evaluation with Offline Data Informed Behavior Policy Design

2023-01-31 · Shuze Liu, Shangtong Zhang

Most reinforcement learning practitioners evaluate their policies with online Monte Carlo estimators for either hyperparameter tuning or testing different algorithmic design choices, where the policy is repeatedly executed in the environment to get the average outcome. Such massive interactions with the environment are prohibitive in many scenarios. In this paper, we propose novel methods that improve the data efficiency of online Monte Carlo estimators while maintaining their unbiasedness. We first propose a tailored closed-form behavior policy that provably reduces the variance of an online Monte Carlo estimator. We then design efficient algorithms to learn this closed-form behavior policy from previously collected offline data. Theoretical analysis is provided to characterize how the behavior policy learning error affects the amount of reduced variance. Compared with previous works, our method achieves better empirical performance in a broader set of environments, with fewer requirements for offline data.

📄 PDF Abstract BibTeX arXiv:2301.13734

Code (1)

shuzeliu/behavior-policy-design-for-policy-evaluation 공식 구현 pytorch

Tasks

Management

Similar Papers 제목 키워드 기반

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning

2024-12-12 · Prajwal Koirala, Zhanhong Jiang, Soumik Sarkar, Cody Fleming

Safe offline reinforcement learning aims to learn policies that maximize cumulative rewards while adhering to safety constraints, using only offline data for training. A key challenge is balancing safety and performance,…

Efficient Online Learning with Offline Datasets for Infinite Horizon MDPs: A Bayesian Approach

2023-10-17 · Dengwang Tang, Rahul Jain, Botao Hao, Zheng Wen

In this paper, we study the problem of efficient online reinforcement learning in the infinite horizon setting when there is an offline dataset to start with. We assume that the offline dataset is generated by an expert …

Imitation Learning

Matrix Estimation for Offline Reinforcement Learning with Low-Rank Structure

2023-05-24 · Xumei Xi, Christina Lee Yu, Yudong Chen

We consider offline Reinforcement Learning (RL), where the agent does not interact with the environment and must rely on offline data collected using a behavior policy. Previous works provide policy evaluation guarantees…

Matrix Completionreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Offline Reinforcement Learning with Reverse Model-based Imagination

2021-10-01 · NeurIPS 2021 12 · Jianhao Wang, Wenzhe Li, Haozhe Jiang, Guangxiang Zhu 외

In offline reinforcement learning (offline RL), one of the main challenges is to deal with the distributional shift between the learning policy and the given dataset. To address this problem, recent offline RL methods at…

Data AugmentationmodelOffline RLreinforcement-learning+2

Offline Reinforcement Learning for Warehouse SLAM Throughput Control

2026-06-22 · Tina Dongxu Li, Mouhacine Benosman, Rajat Kumar, Kevin Tan 외 arxiv

We present an offline reinforcement learning (RL) framework for optimizing SLAM throughput control in a warehouse fulfillment environment. SLAM (Scan/Label/Apply/Manifest) throughput directly influences system congestion…

Reinforcement LearningOffline RL