paper-with-me

Papers

Leveraging Offline Data in Online Reinforcement Learning

2022-11-09 · Andrew Wagenmaker, Aldo Pacchiano

Two central paradigms have emerged in the reinforcement learning (RL) community: online RL and offline RL. In the online RL setting, the agent has no prior knowledge of the environment, and must interact with it in order to find an $\epsilon$-optimal policy. In the offline RL setting, the learner instead has access to a fixed dataset to learn from, but is unable to otherwise interact with the environment, and must obtain the best policy it can from this offline data. Practical scenarios often motivate an intermediate setting: if we have some set of offline data and, in addition, may also interact with the environment, how can we best use the offline data to minimize the number of online interactions necessary to learn an $\epsilon$-optimal policy? In this work, we consider this setting, which we call the \textsf{FineTuneRL} setting, for MDPs with linear structure. We characterize the necessary number of online samples needed in this setting given access to some offline dataset, and develop an algorithm, \textsc{FTPedel}, which is provably optimal, up to $H$ factors. We show through an explicit example that combining offline data with online interactions can lead to a provable improvement over either purely offline or purely online RL. Finally, our results illustrate the distinction between \emph{verifiable} learning, the typical setting considered in online RL, and \emph{unverifiable} learning, the setting often considered in offline RL, and show that there is a formal separation between these regimes.

📄 PDF Abstract BibTeX arXiv:2211.04974

Code (0)

등록된 구현이 없습니다.

Tasks

Offline RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

SAMG: State-Action-Aware Offline-to-Online Reinforcement Learning with Offline Model Guidance

2024-10-24 · Liyu Zhang, Haochi Wu, Xu Wan, Quan Kong 외

The offline-to-online (O2O) paradigm in reinforcement learning (RL) utilizes pre-trained models on offline datasets for subsequent online fine-tuning. However, conventional O2O RL algorithms typically require maintaining…

D4RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Efficient Post-training of LLMs for Code Generation With Offline Reinforcement Learning

2026-05-27 · Mingze Wu, Abhinav Anand, Shweta Verma, Mira Mezini arxiv

Post-training using online reinforcement learning (RL) is an important training step for LLMs, including code-generating models. However, online RL for code generation involves LLM inference and verification of the gener…

Reinforcement LearningCode GenerationOffline RL

Efficient Reinforcement Learning by Guiding Generalist World Models with Non-Curated Data

2025-02-26 · Yi Zhao, Aidan Scannell, Wenshuai Zhao, Yuxin Hou 외

Leveraging offline data is a promising way to improve the sample efficiency of online reinforcement learning (RL). This paper expands the pool of usable data for offline-to-online RL by leveraging abundant non-curated da…

Attributereinforcement-learningReinforcement LearningReinforcement Learning (RL)

Reward-agnostic Fine-tuning: Provable Statistical Benefits of Hybrid Reinforcement Learning

2023-05-17 · NeurIPS 2023 11

This paper studies tabular reinforcement learning (RL) in the hybrid setting, which assumes access to both an offline dataset and online interactions with the unknown environment. A central question boils down to how to …

Offline RLreinforcement-learningReinforcement Learning (RL)

Towards Robust Offline-to-Online Reinforcement Learning via Uncertainty and Smoothness

2023-09-29 · Xiaoyu Wen, Xudong Yu, Rui Yang, HaoYuan Chen 외

To obtain a near-optimal policy with fewer interactions in Reinforcement Learning (RL), a promising approach involves the combination of offline RL, which enhances sample efficiency by leveraging offline datasets, and on…

Offline RLreinforcement-learningReinforcement Learning (RL)