Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
We consider a hybrid reinforcement learning setting (Hybrid RL), in which an agent has access to an offline dataset and the ability to collect experience via real-world online interaction. The framework mitigates the challenges that arise in both pure offline and online RL settings, allowing for the design of simple and highly effective algorithms, in both theory and practice. We demonstrate these advantages by adapting the classical Q learning/iteration algorithm to the hybrid setting, which we call Hybrid Q-Learning or Hy-Q. In our theoretical results, we prove that the algorithm is both computationally and statistically efficient whenever the offline dataset supports a high-quality policy and the environment has bounded bilinear rank. Notably, we require no assumptions on the coverage provided by the initial distribution, in contrast with guarantees for policy gradient/iteration methods. In our experimental results, we show that Hy-Q with neural network function approximation outperforms state-of-the-art online, offline, and hybrid RL baselines on challenging benchmarks, including Montezuma's Revenge.
Code (1)
Tasks
Montezuma's RevengeQ-LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A Natural Extension To Online Algorithms For Hybrid RL With Limited Coverage
Hybrid Reinforcement Learning (RL), leveraging both online and offline data, has garnered recent interest, yet research on its provable benefits remains sparse. Additionally, many existing hybrid RL algorithms (Song et a…
Efficient ExplorationReinforcement Learning (RL)Reward-agnostic Fine-tuning: Provable Statistical Benefits of Hybrid Reinforcement Learning
This paper studies tabular reinforcement learning (RL) in the hybrid setting, which assumes access to both an offline dataset and online interactions with the unknown environment. A central question boils down to how to …
Offline RLreinforcement-learningReinforcement Learning (RL)Sample-Mean Anchored Thompson Sampling for Offline-to-Online Learning with Distribution Shift
Offline-to-online learning aims to improve online decision-making by leveraging offline logged data. A central challenge in this setting is the distribution shift between offline and online environments. While some exist…
Hybrid Online-Offline Learning for Task Offloading in Mobile Edge Computing Systems
We consider a multi-user multi-server mobile edge computing (MEC) system, in which users arrive on a network randomly over time and generate computation tasks, which will be computed either locally on their own computing…
Edge-computingLearning from Offline and Online Experiences: A Hybrid Adaptive Operator Selection Framework
In many practical applications, usually, similar optimisation problems or scenarios repeatedly appear. Learning from previous problem-solving experiences can help adjust algorithm components of meta-heuristics, e.g., ada…