paper-with-me

홈 › Papers

Read and Reap the Rewards: Learning to Play Atari with the Help of Instruction Manuals

2023-02-09 · NeurIPS 2023 11 · Yue Wu, Yewen Fan, Paul Pu Liang, Amos Azaria, Yuanzhi Li, Tom M. Mitchell

High sample complexity has long been a challenge for RL. On the other hand, humans learn to perform tasks not only from interaction or demonstrations, but also by reading unstructured text documents, e.g., instruction manuals. Instruction manuals and wiki pages are among the most abundant data that could inform agents of valuable features and policies or task-specific environmental dynamics and reward structures. Therefore, we hypothesize that the ability to utilize human-written instruction manuals to assist learning policies for specific tasks should lead to a more efficient and better-performing agent. We propose the Read and Reward framework. Read and Reward speeds up RL algorithms on Atari games by reading manuals released by the Atari game developers. Our framework consists of a QA Extraction module that extracts and summarizes relevant information from the manual and a Reasoning module that evaluates object-agent interactions based on information from the manual. An auxiliary reward is then provided to a standard A2C RL agent, when interaction is detected. Experimentally, various RL algorithms obtain significant improvement in performance and training speed when assisted by our design.

📄 PDF Abstract BibTeX arXiv:2302.04449

Code (0)

등록된 구현이 없습니다.

Tasks

Atari Games

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
A2C A2C, or Advantage Actor Critic, is a synchronous version of the A3C policy gradient method. As an alternative to the asynchronous…

Similar Papers 제목 키워드 기반

Combining Experience Replay with Exploration by Random Network Distillation

2019-05-18 · Francesco Sovrano

Our work is a simple extension of the paper "Exploration by Random Network Distillation". More in detail, we show how to efficiently combine Intrinsic Rewards with Experience Replay in order to achieve more efficient and…

Atari GamesMontezuma's Revenge

Data-Efficient Exploration with Self Play for Atari

2021-06-13 · ICML Workshop URL 2021 7 · Michael Laskin, Catherine Cang, Ryan Rudes, Pieter Abbeel

Most reinforcement learning (RL) algorithms rely on hand-crafted extrinsic rewards to learn skills. However, crafting a reward function for each skill is not scalable and results in narrow agents that learn reward-specif…

Efficient ExplorationReinforcement Learning (RL)

Playing Atari Games with Deep Reinforcement Learning and Human Checkpoint Replay

2016-07-18 · Ionel-Alexandru Hosu, Traian Rebedea

This paper introduces a novel method for learning how to play the most difficult Atari 2600 games from the Arcade Learning Environment using deep reinforcement learning. The proposed method, human checkpoint replay, cons…

Atari GamesDeep Reinforcement LearningMontezuma's Revengereinforcement-learning+2

Playing Atari with Deep Reinforcement Learning

2013-12-19 · Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves 외

We present the first deep learning model to successfully learn control policies directly from high-dimensional sensory input using reinforcement learning. The model is a convolutional neural network, trained with a varia…

Atari GamesDeep Reinforcement LearningMulti-Goal Reinforcement LearningQ-Learning+2

Learning Actions and Control of Focus of Attention with a Log-Polar-like Sensor

2023-09-22 · Robin Göransson, Volker Krueger

With the long-term goal of reducing the image processing time on an autonomous mobile robot in mind we explore in this paper the use of log-polar like image data with gaze control. The gaze control is not done on the Car…

Atari GamesDeep Reinforcement Learning