paper-with-me

홈 › Papers

Search on the Replay Buffer: Bridging Planning and Reinforcement Learning

2019-06-12 · NeurIPS 2019 12 · Benjamin Eysenbach, Ruslan Salakhutdinov, Sergey Levine

The history of learning for control has been an exciting back and forth between two broad classes of algorithms: planning and reinforcement learning. Planning algorithms effectively reason over long horizons, but assume access to a local policy and distance metric over collision-free paths. Reinforcement learning excels at learning policies and the relative values of states, but fails to plan over long horizons. Despite the successes of each method in various domains, tasks that require reasoning over long horizons with limited feedback and high-dimensional observations remain exceedingly challenging for both planning and reinforcement learning algorithms. Frustratingly, these sorts of tasks are potentially the most useful, as they are simple to design (a human only need to provide an example goal state) and avoid reward shaping, which can bias the agent towards finding a sub-optimal solution. We introduce a general control algorithm that combines the strengths of planning and reinforcement learning to effectively solve these tasks. Our aim is to decompose the task of reaching a distant goal state into a sequence of easier tasks, each of which corresponds to reaching a subgoal. Planning algorithms can automatically find these waypoints, but only if provided with suitable abstractions of the environment -- namely, a graph consisting of nodes and edges. Our main insight is that this graph can be constructed via reinforcement learning, where a goal-conditioned value function provides edge weights, and nodes are taken to be previously seen observations in a replay buffer. Using graph search over our replay buffer, we can automatically generate this sequence of subgoals, even in image-based environments. Our algorithm, search on the replay buffer (SoRB), enables agents to solve sparse reward tasks over one hundred steps, and generalizes substantially better than standard RL algorithms.

📄 PDF Abstract BibTeX arXiv:1906.05253

Code (1)

google-research/google-research 공식 구현 tf

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Fast deep reinforcement learning using online adjustments from the past

2018-10-18 · NeurIPS 2018 12 · Steven Hansen, Pablo Sprechmann, Alexander Pritzel, André Barreto 외

We propose Ephemeral Value Adjusments (EVA): a means of allowing deep reinforcement learning agents to rapidly adapt to experience in their replay buffer. EVA shifts the value predicted by a neural network with an estima…

Atari GamesDeep Reinforcement Learningreinforcement-learningReinforcement Learning+2

Analysis of Stochastic Processes through Replay Buffers

2022-06-26 · Shirli Di Castro Shashua, Shie Mannor, Dotan Di-Castro

Replay buffers are a key component in many reinforcement learning schemes. Yet, their theoretical properties are not fully understood. In this paper we analyze a system where a stochastic process X is pushed into a repla…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Replay Buffer with Local Forgetting for Adapting to Local Environment Changes in Deep Model-Based Reinforcement Learning

2023-03-15 · Ali Rahimi-Kalahroudi, Janarthanan Rajendran, Ida Momennejad, Harm van Seijen 외

One of the key behavioral characteristics used in neuroscience to determine whether the subject of study -- be it a rodent or a human -- exhibits model-based learning is effective adaptation to local changes in the envir…

Model-based Reinforcement Learningreinforcement-learningReinforcement Learning (RL)

B2RL: An open-source Dataset for Building Batch Reinforcement Learning

2022-09-30 · Hsin-Yu Liu, Xiaohan Fu, Bharathan Balaji, Rajesh Gupta 외

Batch reinforcement learning (BRL) is an emerging research area in the RL community. It learns exclusively from static datasets (i.e. replay buffers) without interaction with the environment. In the offline settings, exi…

Managementreinforcement-learningReinforcement LearningReinforcement Learning (RL)

ARROW: Augmented Replay for RObust World models

2026-03-12 · Abdulaziz Alyahya, Abdallah Al Siyabi, Markus R. Ernst, Luke Yang 외 arxiv

Continual reinforcement learning challenges agents to acquire new skills while retaining previously learned ones with the goal of improving performance in both past and future tasks. Most existing approaches rely on mode…

Reinforcement Learning