paper-with-me

Papers

Offline Reinforcement Learning Under Value and Density-Ratio Realizability: The Power of Gaps

2022-03-25 · Jinglin Chen, Nan Jiang

We consider a challenging theoretical problem in offline reinforcement learning (RL): obtaining sample-efficiency guarantees with a dataset lacking sufficient coverage, under only realizability-type assumptions for the function approximators. While the existing theory has addressed learning under realizability and under non-exploratory data separately, no work has been able to address both simultaneously (except for a concurrent work which we compare in detail). Under an additional gap assumption, we provide guarantees to a simple pessimistic algorithm based on a version space formed by marginalized importance sampling (MIS), and the guarantee only requires the data to cover the optimal policy and the function classes to realize the optimal value and density-ratio functions. While similar gap assumptions have been used in other areas of RL theory, our work is the first to identify the utility and the novel mechanism of gap assumptions in offline RL with weak function approximation.

📄 PDF Abstract BibTeX arXiv:2203.13935

Code (0)

등록된 구현이 없습니다.

Tasks

Offline RLReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Harnessing Density Ratios for Online Reinforcement Learning

2024-01-18 · Philip Amortila, Dylan J. Foster, Nan Jiang, Ayush Sekhari 외

The theories of offline and online reinforcement learning, despite having evolved in parallel, have begun to show signs of the possibility for a unification, with algorithms and analysis techniques for one setting often …

Offline RLreinforcement-learningReinforcement Learning

Constrained Policy Optimization with Explicit Behavior Density for Offline Reinforcement Learning

2023-01-28 · NeurIPS 2023 11 · Jing Zhang, Chi Zhang, Wenjia Wang, Bing-Yi Jing

Due to the inability to interact with the environment, offline reinforcement learning (RL) methods face the challenge of estimating the Out-of-Distribution (OOD) points. Existing methods for addressing this issue either …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Offline Reinforcement Learning with Soft Behavior Regularization

2021-10-14 · Haoran Xu, Xianyuan Zhan, Jianxiong Li, Honglei Yin

Most prior approaches to offline reinforcement learning (RL) utilize \textit{behavior regularization}, typically augmenting existing off-policy actor critic algorithms with a penalty measuring divergence between the poli…

continuous-controlContinuous Controlreinforcement-learningReinforcement Learning+1

Offline Oracle-Efficient Learning for Contextual MDPs via Layerwise Exploration-Exploitation Tradeoff

2024-05-28 · Jian Qian, Haichen Hu, David Simchi-Levi

Motivated by the recent discovery of a statistical and computational reduction from contextual bandits to offline regression (Simchi-Levi and Xu, 2021), we address the general (stochastic) Contextual Markov Decision Proc…

Density EstimationMulti-Armed Bandits

A Simple Unified Uncertainty-Guided Framework for Offline-to-Online Reinforcement Learning

2023-06-13 · Siyuan Guo, Yanchao Sun, Jifeng Hu, Sili Huang 외

Offline reinforcement learning (RL) provides a promising solution to learning an agent fully relying on a data-driven paradigm. However, constrained by the limited quality of the offline dataset, its performance is often…

D4RLEfficient ExplorationOffline RLreinforcement-learning+1