paper-with-me

홈 › Papers

Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

2022-10-13 · Yuda Song, Yifei Zhou, Ayush Sekhari, J. Andrew Bagnell, Akshay Krishnamurthy, Wen Sun

We consider a hybrid reinforcement learning setting (Hybrid RL), in which an agent has access to an offline dataset and the ability to collect experience via real-world online interaction. The framework mitigates the challenges that arise in both pure offline and online RL settings, allowing for the design of simple and highly effective algorithms, in both theory and practice. We demonstrate these advantages by adapting the classical Q learning/iteration algorithm to the hybrid setting, which we call Hybrid Q-Learning or Hy-Q. In our theoretical results, we prove that the algorithm is both computationally and statistically efficient whenever the offline dataset supports a high-quality policy and the environment has bounded bilinear rank. Notably, we require no assumptions on the coverage provided by the initial distribution, in contrast with guarantees for policy gradient/iteration methods. In our experimental results, we show that Hy-Q with neural network function approximation outperforms state-of-the-art online, offline, and hybrid RL baselines on challenging benchmarks, including Montezuma's Revenge.

📄 PDF Abstract BibTeX arXiv:2210.06718

Code (1)

yudasong/hyq 공식 구현 pytorch

Tasks

Montezuma's RevengeQ-Learning

Methods 이 논문이 사용한 방법론

Q-Learning Q-Learning is an off-policy temporal difference control algorithm: $$Q\left(S\_{t}, A\_{t}\right) \leftarrow Q\left(S\_{t}, A\_{t}\right) + \alpha\left[R_{t+1} +…

Similar Papers 제목 키워드 기반

A Natural Extension To Online Algorithms For Hybrid RL With Limited Coverage

2024-03-07 · Kevin Tan, Ziping Xu

Hybrid Reinforcement Learning (RL), leveraging both online and offline data, has garnered recent interest, yet research on its provable benefits remains sparse. Additionally, many existing hybrid RL algorithms (Song et a…

Efficient ExplorationReinforcement Learning (RL)

Reward-agnostic Fine-tuning: Provable Statistical Benefits of Hybrid Reinforcement Learning

2023-05-17 · NeurIPS 2023 11

This paper studies tabular reinforcement learning (RL) in the hybrid setting, which assumes access to both an offline dataset and online interactions with the unknown environment. A central question boils down to how to …

Offline RLreinforcement-learningReinforcement Learning (RL)

Sample-Mean Anchored Thompson Sampling for Offline-to-Online Learning with Distribution Shift

2026-05-11 · Bochao Li, Yao Fu, Wei Chen, Fang Kong arxiv

Offline-to-online learning aims to improve online decision-making by leveraging offline logged data. A central challenge in this setting is the distribution shift between offline and online environments. While some exist…

Hybrid Online-Offline Learning for Task Offloading in Mobile Edge Computing Systems

2024-02-19 · Muhammad Sohaib, Sang-Woon Jeon, Wei Yu

We consider a multi-user multi-server mobile edge computing (MEC) system, in which users arrive on a network randomly over time and generate computation tasks, which will be computed either locally on their own computing…

Edge-computing

Learning from Offline and Online Experiences: A Hybrid Adaptive Operator Selection Framework

2024-04-16 · Jiyuan Pei, Jialin Liu, Yi Mei

In many practical applications, usually, similar optimisation problems or scenarios repeatedly appear. Learning from previous problem-solving experiences can help adjust algorithm components of meta-heuristics, e.g., ada…